论文解读

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs

语音识别 | 7/10

 · 更新于 2026-09-24 · 约 20 分钟 · 9935 字 阅读 →
论文解读

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

语音识别 | 7.3/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8554 字 阅读 →
论文解读

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

语音交互 | 9.2/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8435 字 阅读 →
论文解读

WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS

语音合成 | 7.2/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7351 字 阅读 →
论文解读

Reinforcement Learning for Data-Efficient Code-Switched ASR

语音识别 | 5.3/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5883 字 阅读 →
论文解读

Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization

语音编码 | 7.9/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5637 字 阅读 →
论文解读

\(\tau\)-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains

语音交互 | 9.1/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8721 字 阅读 →
论文解读

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

语音合成 | 6.6/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6527 字 阅读 →
论文解读

LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues

语音交互 | 8.1/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7266 字 阅读 →
论文解读

Unlocking Speech–Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

语音交互 | 6.7/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7653 字 阅读 →
论文解读

NAVER LABS Europe Submission to the Instruction-following 2026 Short Track

语音翻译 | 6.2/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7251 字 阅读 →
论文解读

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

语音交互 | 8.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5306 字 阅读 →
论文解读

Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs

语音识别 | 9.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5833 字 阅读 →
论文解读

Codec-Robust Attacks on Audio LLMs

音频安全 | 8.3/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8462 字 阅读 →
论文解读

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

语音识别 | 9.3/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7818 字 阅读 →
论文解读

Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models

语音对话系统 | 7.6/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7848 字 阅读 →
论文解读

FormalASR: End-to-End Spoken Chinese to Formal Text

语音识别 | 6/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7432 字 阅读 →
论文解读

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

语音识别 | 6.8/10

 · 更新于 2026-09-24 · 约 22 分钟 · 10575 字 阅读 →
论文解读

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities

音频问答 | 6.5/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8983 字 阅读 →
论文解读

From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents

模型评估 | 6.3/10

 · 更新于 2026-09-24 · 约 22 分钟 · 10971 字 阅读 →