论文解读

AudioX: A Unified Framework for Anything-to-Audio Generation

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10720 字 阅读 →
论文解读

AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4850 字 阅读 →
论文解读

AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration

视频描述生成 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4718 字 阅读 →
论文解读

Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4990 字 阅读 →
论文解读

Can Vision-Language Models Answer Face to Face Questions in the Real-World?

音频问答 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4402 字 阅读 →
论文解读

CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval

音频检索 音乐理解 | 6.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5689 字 阅读 →
论文解读

Data-Centric Lessons To Improve Speech-Language Pretraining

语音问答 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4090 字 阅读 →
论文解读

DiffSDA: Unsupervised Diffusion Sequential Disentanglement Across Modalities

序列解耦 | 8.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6401 字 阅读 →
论文解读

DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations

语音对话系统 | 9.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4530 字 阅读 →
论文解读

Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning

音频问答 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4607 字 阅读 →
论文解读

Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention

语音分离 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5777 字 阅读 →
论文解读

End-to-end Listen, Look, Speak and Act

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5323 字 阅读 →
论文解读

Entropy-Monitored Kernelized Token Distillation for Audio-Visual Compression

音视频事件检测 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5068 字 阅读 →
论文解读

FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows

跨模态生成 | 9.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5159 字 阅读 →
论文解读

From Natural Alignment to Conditional Controllability in Multimodal Dialogue

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4836 字 阅读 →
论文解读

From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5679 字 阅读 →
论文解读

GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models

音乐理解 | 7.0/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3318 字 阅读 →
论文解读

Group Cognition Learning: Making Everything Better Through Governed Two-Stage Agents Collaboration

多模态模型 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5604 字 阅读 →
论文解读

Human Behavior Atlas: Benchmarking Unified Psychological And Social Behavior Understanding

多模态模型 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5063 字 阅读 →
论文解读

Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4701 字 阅读 →