论文解读

Closing the Gap Between Text and Speech Understanding in LLMs

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8119 字 阅读 →
论文解读

Deep Learning with Learnable Product-Structured Activations

音频分类 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5033 字 阅读 →
论文解读

DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5139 字 阅读 →
论文解读

Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention

语音分离 | 9.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5833 字 阅读 →
论文解读

End-to-end Listen, Look, Speak and Act

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6653 字 阅读 →
论文解读

From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Training

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5535 字 阅读 →
论文解读

Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech Synthesis

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6858 字 阅读 →
论文解读

Learnable Fractional Superlets with a Spectro-Temporal Emotion Encoder for Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4089 字 阅读 →
论文解读

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5987 字 阅读 →
论文解读

RoboOmni: Proactive Robot Manipulation in Omni-modal Context

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4469 字 阅读 →
论文解读

TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling

语音对话系统 | 8.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3884 字 阅读 →
论文解读

Towards True Speech-to-Speech Models Without Text Guidance

语音对话系统 | 9.1/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4831 字 阅读 →
论文解读

TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization

语音转换 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6178 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-02

共分析 4 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 15 分钟 · 7026 字 阅读 →
论文解读

Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge

语音对话系统 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3578 字 阅读 →
论文解读

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6664 字 阅读 →
论文解读

Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device

语音生物标志物 | 7.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6353 字 阅读 →
论文解读

Text-Utilization for Encoder-dominated Speech Recognition Models

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 6 分钟 · 2856 字 阅读 →
论文解读

A Speech-Driven Paradigm for Physics-Informed Modeling of Coupled Micro-Speakers

音频生成 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4668 字 阅读 →
论文解读

Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5056 字 阅读 →