每日研究速递

语音/音乐/音频论文速递 2026-04-25

共分析 2 篇语音/AI 论文

 · 更新于 2026-09-10 · 约 6 分钟 · 2986 字 阅读 →
论文解读

"This Wasn't Made for Me": Recentering User Experience and Emotional Impact in the Evaluation of ASR Bias

语音识别 | 7.0/10

 · 更新于 2026-09-10 · 约 4 分钟 · 1997 字 阅读 →
论文解读

ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis

语音合成 | 7.0/10

 · 更新于 2026-09-10 · 约 14 分钟 · 6680 字 阅读 →
论文解读

AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA

音频问答 | 6.5/10

 · 更新于 2026-09-10 · 约 5 分钟 · 2444 字 阅读 →
论文解读

Beyond Rules: Towards Basso Continuo Personal Style Identification

音乐理解 | 7.0/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3212 字 阅读 →
论文解读

DiariZen Explained: A Tutorial for the Open Source State-of-the-Art Speaker Diarization Pipeline

说话人分离 | 6.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4881 字 阅读 →
论文解读

Dilated CNNs for Periodic Signal Processing: A Low-Complexity Approach

语音增强 | 6.5/10

 · 更新于 2026-09-10 · 约 5 分钟 · 2311 字 阅读 →
论文解读

Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4532 字 阅读 →
论文解读

Evaluation of Automatic Speech Recognition Using Generative Large Language Models

语音识别 | 7.5/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3189 字 阅读 →
论文解读

Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge

语音对话系统 | 6.5/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3206 字 阅读 →
论文解读

Hierarchical Policy Optimization for Simultaneous Translation of Unbounded Speech

语音翻译 | 7.5/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3626 字 阅读 →
论文解读

Low-Rank Adaptation Redux for Large Models

大语言模型 | 5.5/10

 · 更新于 2026-09-10 · 约 4 分钟 · 1870 字 阅读 →
论文解读

MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

语音合成 | 7.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4902 字 阅读 →
论文解读

Materialistic RIR: Material Conditioned Realistic RIR Generation

音频生成 | 7.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5449 字 阅读 →
论文解读

MER 2026: From Discriminative Emotion Recognition to Generative Emotion Understanding

语音情感识别 | 6.0/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5670 字 阅读 →
论文解读

Misinformation Span Detection in Videos via Audio Transcripts

音频安全 | 7.5/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4016 字 阅读 →
论文解读

Phonological Subspace Collapse Is Aetiology-Specific and Cross-Lingually Stable: Evidence from 3,374 Speakers

Phonological Subspace Collapse Is Aetiology-Specific and Cross-Lingually Stable: Evidence from 3,374 Speakers

 · 更新于 2026-09-10 · 约 1 分钟 · 30 字 阅读 →
论文解读

Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

语音合成 | 7.5/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3910 字 阅读 →
论文解读

Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5231 字 阅读 →
论文解读

Sema: Semantic Transport for Real-Time Multimodal Agents

实时处理 | 6.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4888 字 阅读 →