论文解读

DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations

语音对话系统 | 9.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4530 字 阅读 →
论文解读

Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning

音频问答 | 8.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4607 字 阅读 →
论文解读

EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Models

基准测试 | 7.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4360 字 阅读 →
论文解读

Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention

语音分离 | 7.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5777 字 阅读 →
论文解读

EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

语音情感识别 | 8.0/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3695 字 阅读 →
论文解读

End-to-end Listen, Look, Speak and Act

语音对话系统 | 8.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5323 字 阅读 →
论文解读

Entropy-Monitored Kernelized Token Distillation for Audio-Visual Compression

音视频事件检测 | 8.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5068 字 阅读 →
论文解读

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation

音频生成 | 7.5/10

 · 更新于 2026-09-10 · 约 14 分钟 · 6971 字 阅读 →
论文解读

FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates

语音合成 | 9.0/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5627 字 阅读 →
论文解读

FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions

语音合成 | 8.0/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5075 字 阅读 →
论文解读

Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation

音频生成 | 8.0/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6318 字 阅读 →
论文解读

FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows

跨模态生成 | 9.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5159 字 阅读 →
论文解读

From Birdsong to Rumbles: Classifying Elephant Calls with Out-of-Species Embeddings

音频分类 | 6.5/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6394 字 阅读 →
论文解读

From Natural Alignment to Conditional Controllability in Multimodal Dialogue

语音合成 | 8.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4836 字 阅读 →
论文解读

From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training

语音对话系统 | 8.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5679 字 阅读 →
论文解读

GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models

音乐理解 | 7.0/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3318 字 阅读 →
论文解读

Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction

音乐生成 | 7.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4702 字 阅读 →
论文解读

Gogo: Group-wise granularity-ordered codec for stable and efficient speech generation

语音合成 | 7.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5137 字 阅读 →
论文解读

Group Cognition Learning: Making Everything Better Through Governed Two-Stage Agents Collaboration

多模态模型 | 7.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5604 字 阅读 →
论文解读

Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech Synthesis

语音合成 | 8.0/10

 · 更新于 2026-09-10 · 约 20 分钟 · 9658 字 阅读 →