论文解读

Temporal-Spatial Decouple Before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis

情感分析 | 7.5/10

 · 更新于 2026-09-16 · 约 13 分钟 · 6191 字 阅读 →
论文解读

Temporally Heterogeneous Graph Contrastive Learning for Multimodal Acoustic Event Classification

音频事件检测 | 8.5/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3902 字 阅读 →
论文解读

Test Time Adaptation for Speech Emotion Recognition

语音情感识别 | 7.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4724 字 阅读 →
论文解读

Test-Time Scaling for Auditory Cognition in Audio Language Models

音频问答 | 7.0/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4888 字 阅读 →
论文解读

Testing The Efficient Coding Hypothesis Beyond Humans: The Auditory Kernels of Bat Vocalizations

生物声学 | 7.5/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3800 字 阅读 →
论文解读

Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment

音乐生成 | 7.5/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3668 字 阅读 →
论文解读

Text2Move: Text-To-Moving Sound Generation via Trajectory Prediction and Temporal Alignment

空间音频 | 8.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4225 字 阅读 →
论文解读

TextlessRAG: End-to-End Visual Document RAG by Speech without Text

语音问答 | 8.5/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4218 字 阅读 →
论文解读

The 3rd Clarity Prediction Challenge: A Machine Learning Challenge for Hearing aid Speech Intelligibility Prediction

语音增强 | 7.5/10

 · 更新于 2026-09-16 · 约 7 分钟 · 3264 字 阅读 →
论文解读

The Curious Case of Visual Grounding: Different Effects for Speech-and Text-Based Language Encoders

模型评估 | 8.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4508 字 阅读 →
论文解读

The Impact of Audio Watermarking on Audio Anti-Spoofing Countermeasures

音频深度伪造检测 | 8.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4839 字 阅读 →
论文解读

The Muse Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMs

音乐理解 | 8.5/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3655 字 阅读 →
论文解读

The Role of Prosodic and Lexical Cues in Turn-Taking with Self-Supervised Speech Representations

语音对话系统 | 7.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4558 字 阅读 →
论文解读

The Singing Voice Conversion Challenge 2025: From Singer Identity Conversion to Singing Style Conversion

歌唱语音转换 | 7.0/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3758 字 阅读 →
论文解读

The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models

基准测试 | 7.0/10

 · 更新于 2026-09-16 · 约 8 分钟 · 3637 字 阅读 →
论文解读

The Synergistic Role of Audio and Large Video-Language Model in Source-Free Video Domain Adaptation

领域适应 | 7.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4434 字 阅读 →
论文解读

Theory and Application of Circular Relative Harmonic Coefficients

声源定位 | 7.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4582 字 阅读 →
论文解读

Thinking While Listening: Simple Test Time Scaling for Audio Classification

音频分类 | 6.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4620 字 阅读 →
论文解读

Three Seconds is Sufficient: A Multi-Pronged Framework for Model-Based Speaker Adaptation in ASR Under Data-Scarce Conditions

语音识别 | 7.0/10

 · 更新于 2026-09-16 · 约 9 分钟 · 4318 字 阅读 →
论文解读

TICL: Text-Embedding KNN for Speech in-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models

语音识别 | 7.5/10

 · 更新于 2026-09-16 · 约 10 分钟 · 4710 字 阅读 →