论文解读

Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and Algorithm

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10471 字 阅读 →
论文解读

USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding

音频编码 | 9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 6009 字 阅读 →
论文解读

Logit Distillation on Manifolds: Mapping by Learning

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6262 字 阅读 →
论文解读

Raon-Speech Technical Report

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5464 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-30

共分析 6 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 17 分钟 · 8127 字 阅读 →
论文解读

Data-Efficient On-Policy Distillation for Automatic Speech Recognition

语音识别 | 5.1/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4830 字 阅读 →
论文解读

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

语音合成 | 8.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6179 字 阅读 →
论文解读

State-Anchored Complete-View Distillation for Robust Conversational Multimodal Emotion Recognition

语音情感识别 | 8/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7970 字 阅读 →
论文解读

OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation

音频检索 | 9.2/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5784 字 阅读 →
论文解读

S2Accompanist: A Semantic-Aware and Structure-Guided Diffusion Model for Music Accompaniment Generation

音乐生成 | 5.6/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9578 字 阅读 →
论文解读

Sometin Beta Pass Notin (SBPN): Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation

语音识别 | 6.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7571 字 阅读 →
论文解读

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling

音频编码 | 7.0/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8374 字 阅读 →
论文解读

Evaluating the Expressive Appropriateness of Speech in Rich Contexts

语音质量评估 | 7.2/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9625 字 阅读 →
论文解读

Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation

语音增强 | 7.2/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10260 字 阅读 →
论文解读

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

语音大模型 | 8.0/10

 · 更新于 2026-09-25 · 约 29 分钟 · 14118 字 阅读 →
论文解读

Modality-Aware Contrastive and Uncertainty-Regularized Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7828 字 阅读 →
论文解读

To Fuse or to Drop? Dual-Path Learning for Resolving Modality Conflicts in Multimodal Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6769 字 阅读 →
论文解读

AsymK-Talker: Real-Time and Long-Horizon Talking Head Generation via Asymmetric Kernel Distillation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6222 字 阅读 →
论文解读

Private Speech Classification without Collapse: Stabilized DP Training and Offline Distillation

音频分类 | 6.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5535 字 阅读 →
论文解读

Closing the Gap Between Text and Speech Understanding in LLMs

语音大模型 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4518 字 阅读 →