论文解读

OpenSTBench: Beyond Semantic Evaluation for Speech Translation

语音翻译 | 6.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6238 字 阅读 →
论文解读

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue

语音合成 | 8.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8091 字 阅读 →
论文解读

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

语音合成 | 10/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7690 字 阅读 →
论文解读

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion

语音合成 | 8.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7971 字 阅读 →
论文解读

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5495 字 阅读 →
论文解读

Raon-Speech Technical Report

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5464 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-30

共分析 6 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 17 分钟 · 8127 字 阅读 →
论文解读

Dial HEALTHDIAL for Advice: A Multilingual and Multi-Parallel Spoken Dialogue Dataset for Knowledge-Grounded Information Seeking

语音合成 | 8.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5895 字 阅读 →
论文解读

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

语音合成 | 8.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6179 字 阅读 →
论文解读

MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables

语音合成 | 7.3/10

 · 更新于 2026-09-25 · 约 3 分钟 · 1272 字 阅读 →
论文解读

OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants

语音识别 | 7.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6454 字 阅读 →
论文解读

The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models

语音识别 | 7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6823 字 阅读 →
论文解读

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents

语音合成 | 7.9/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7179 字 阅读 →
论文解读

Benchmarking AI for low-resource contexts: Thinking beyond leaderboards

语音识别 | 5.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6151 字 阅读 →
论文解读

Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios

语音合成 | 9.9/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7835 字 阅读 →
论文解读

Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text

音频生成 | 8.6/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7601 字 阅读 →
论文解读

I Hear, Therefore I Trust: A Socio-Technical Investigation of Humans as Synthetic Speech Detectors

语音合成 | 6.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6689 字 阅读 →
论文解读

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

语音合成 | 8/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4782 字 阅读 →
论文解读

Can We Hear from Events? Generating Speech from Event Camera

语音合成 | 7.8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6885 字 阅读 →
论文解读

Continual Speaker Identity Unlearning with Minimal Interference

语音合成 | 8.3/10

 · 更新于 2026-09-25 · 约 4 分钟 · 1998 字 阅读 →