论文解读

Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning

音频问答 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4607 字 阅读 →
论文解读

EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3695 字 阅读 →
论文解读

FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5075 字 阅读 →
论文解读

GaMMA: Towards Joint Global-Temporal Music Understanding in Large Multimodal Models

音乐理解 | 7.0/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3318 字 阅读 →
论文解读

Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction

音乐生成 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4702 字 阅读 →
论文解读

Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards

音频问答 | 8.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4180 字 阅读 →
论文解读

MARS-Sep: Multimodal-Aligned Reinforced Sound Separation

语音分离 | 7.5/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12209 字 阅读 →
论文解读

Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models

音频问答 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4699 字 阅读 →
论文解读

Music Flamingo: Scaling Music Understanding in Audio Language Models

音乐理解 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5841 字 阅读 →
论文解读

Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences

基准测试 数据集 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4749 字 阅读 →
论文解读

ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction

语音对话系统 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4678 字 阅读 →
论文解读

PrismAudio: Decomposed Chain-of-Thought and Multi-dimensional Rewards for Video-to-Audio Generation

音频生成 | 9.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4846 字 阅读 →
论文解读

SpeechJudge: Towards Human-Level Judgment for Speech Naturalness

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8929 字 阅读 →
论文解读

VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation

语音情感识别 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5385 字 阅读 →
论文解读

AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration

音视频 | 8.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6017 字 阅读 →
论文解读

Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning

音频问答 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5380 字 阅读 →
论文解读

EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5631 字 阅读 →
论文解读

Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction

音乐生成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4687 字 阅读 →
论文解读

Human Behavior Atlas: Benchmarking Unified Psychological And Social Behavior Understanding

多模态模型 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5900 字 阅读 →
论文解读

Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards

音频问答 | 8.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4396 字 阅读 →