论文解读

Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention

语音分离 | 9.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5833 字 阅读 →
论文解读

EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

语音情感识别 | 8.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5631 字 阅读 →
论文解读

End-to-end Listen, Look, Speak and Act

语音对话系统 | 8.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6653 字 阅读 →
论文解读

Entropy-Monitored Kernelized Token Distillation for Audio-Visual Compression

音视频 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4449 字 阅读 →
论文解读

FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates

语音合成 | 8.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6115 字 阅读 →
论文解读

FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions

语音合成 | 8.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4459 字 阅读 →
论文解读

Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation

音频生成 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4234 字 阅读 →
论文解读

FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows

音频生成 | 8.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6079 字 阅读 →
论文解读

From Natural Alignment to Conditional Controllability in Multimodal Dialogue

语音合成 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4689 字 阅读 →
论文解读

From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training

语音对话系统 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5535 字 阅读 →
论文解读

Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction

音乐生成 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4687 字 阅读 →
论文解读

Gogo: Group-wise granularity-ordered codec for stable and efficient speech generation

语音合成 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4704 字 阅读 →
论文解读

Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech Synthesis

语音合成 | 8.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6858 字 阅读 →
论文解读

Human Behavior Atlas: Benchmarking Unified Psychological And Social Behavior Understanding

多模态模型 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5900 字 阅读 →
论文解读

Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction

语音对话系统 | 9.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4643 字 阅读 →
论文解读

Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards

音频问答 | 8.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4396 字 阅读 →
论文解读

Instilling an Active Mind in Avatars via Cognitive Simulation

数字人生成 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4161 字 阅读 →
论文解读

InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

视频生成 | 9.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4843 字 阅读 →
论文解读

JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models

音频安全 | 8.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5487 字 阅读 →
论文解读

JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization

音频生成 | 8.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5924 字 阅读 →