论文解读

VoCodec: A Low-bitrate Streamable Neural Speech Codec with Voicing-driven Quantization

语音编码 | 7.2/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5601 字 阅读 →
论文解读

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

语音合成 | 10/10

 · 更新于 2026-09-25 · 约 6 分钟 · 2923 字 阅读 →
论文解读

DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs

语音翻译 | 7.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5915 字 阅读 →
论文解读

Learning When to Think While Listening in Large Audio-Language Models

语音识别 | 8.9/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2186 字 阅读 →
论文解读

OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding

视频理解 | 7.3/10

 · 更新于 2026-09-25 · 约 18 分钟 · 9001 字 阅读 →
论文解读

Contextual Biasing for Streaming ASR via CTC-based Word Spotting

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7082 字 阅读 →
论文解读

Streaming Speech-to-Text Translation with a SpeechLLM

语音翻译 | 6.8/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10616 字 阅读 →
论文解读

Chunkwise Aligners for Streaming Speech Recognition

语音识别 | 6.3/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9896 字 阅读 →
论文解读

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue

语音对话系统 | 6.0/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8653 字 阅读 →
论文解读

Online Segmented Beamforming via Dynamic Programming

声源定位 | 6.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6970 字 阅读 →
论文解读

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7885 字 阅读 →
论文解读

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 26 分钟 · 12846 字 阅读 →
论文解读

End-to-end Listen, Look, Speak and Act

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5323 字 阅读 →
论文解读

MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control

语音合成 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4034 字 阅读 →
论文解读

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4175 字 阅读 →
论文解读

STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4500 字 阅读 →
论文解读

Towards True Speech-to-Speech Models Without Text Guidance

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5080 字 阅读 →
论文解读

TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization

语音转换 语音匿名化 | 8.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6522 字 阅读 →
论文解读

Can Speech LLMs Think while Listening?

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4627 字 阅读 →
论文解读

Can Vision-Language Models Answer Face to Face Questions in the Real-World?

音频问答 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4693 字 阅读 →