论文解读

Improving multichannel speech enhancement through accurate room-acoustic simulations

语音增强 | 6.8/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5624 字 阅读 →
论文解读

Linguistic Bias Mitigation for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck

自监督学习 | 8.6/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5540 字 阅读 →
论文解读

Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection

数据增强 | 7/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6468 字 阅读 →
论文解读

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

语音合成 | 7.3/10

 · 更新于 2026-09-24 · 约 5 分钟 · 2255 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-01

共分析 35 篇语音/AI 论文

 · 更新于 2026-09-24 · 约 102 分钟 · 50963 字 阅读 →
论文解读

AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification

说话人识别 | 7.8/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6161 字 阅读 →
论文解读

How to Leverage Synthetic Speech for LLM-Based ASR Systems?

语音识别 | 8.7/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7623 字 阅读 →
论文解读

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

音乐生成 | 9.4/10

 · 更新于 2026-09-24 · 约 4 分钟 · 1658 字 阅读 →
论文解读

Proteus: Automated Adversarial Robustness Testing for Audio Deepfake Detectors

数据增强 | 5.3/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4768 字 阅读 →
论文解读

Semi-Supervised Sound Event Detection with Conditional Mixup and Embedding-Level Contrastive Loss

对比学习 | 7.6/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5962 字 阅读 →
论文解读

SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation

语音合成 | 6.6/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6319 字 阅读 →
论文解读

Two kinds of robustness are not the same: disentangling fault tolerance and low-SNR robustness in multi-domain event detection on real data

音频事件检测 | 8.9/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4694 字 阅读 →
论文解读

Advancing Speaker-Based Vocal Effort Classification with WavLM and Data Augmentation in Naturalistic Non-Calibrated Speech Recordings

语音增强 | 6.8/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4639 字 阅读 →
论文解读

Do Speech Emphasis Models Generalize across Languages and Emotions?

语音识别 | 7/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4935 字 阅读 →
论文解读

From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection

数据增强 | 8.3/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5898 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-29

共分析 16 篇语音/AI 论文

 · 更新于 2026-09-24 · 约 46 分钟 · 22697 字 阅读 →
论文解读

Listening Like a Judge: A Music-Aware Framework for Automatic Singing Performance Evaluation

歌唱评估 | 8.8/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7529 字 阅读 →
论文解读

Low Resource Multimodal Translation of Nepali Spoken Words into Emotion-Conditioned Sign Language Avatars

语音识别 | 5.3/10

 · 更新于 2026-09-24 · 约 23 分钟 · 11134 字 阅读 →
论文解读

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations

语音合成 | 5.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5965 字 阅读 →
论文解读

Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant

语音识别 | 6.2/10

 · 更新于 2026-09-24 · 约 25 分钟 · 12144 字 阅读 →