论文解读

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy

多模态模型 | 7.4/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5909 字 阅读 →
论文解读

Room for Error: Large-Scale Simulation of Over-the-Air Acoustic Attacks

语音识别 | 6.2/10

 · 更新于 2026-09-07 · 约 16 分钟 · 7709 字 阅读 →
论文解读

Screening Matters: A Comparative Study of Conventional and Crowdsourced Listening Tests

语音质量评估 | 8.4/10

 · 更新于 2026-09-07 · 约 24 分钟 · 11946 字 阅读 →
论文解读

What Was That Again? Certified Robustness for Automatic Speech Recognition

What Was That Again? Certified Robustness for Automatic Speech Recognition

 · 更新于 2026-09-07 · 约 24 分钟 · 11766 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-29

共分析 16 篇语音/AI 论文

 · 更新于 2026-09-07 · 约 46 分钟 · 22697 字 阅读 →
论文解读

A Large-Scale Database and Predictive Model of Listener-Rated Ease of Speech Understanding in Commercial Hearing Aids

语音质量评估 | 8.1/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5406 字 阅读 →
论文解读

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

语音合成 | 6/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4968 字 阅读 →
论文解读

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents

音频分离 | 7.7/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5663 字 阅读 →
论文解读

DNSMOS-C: Improving End-to-end Speech Quality Models via Contrastive Learning

语音质量评估 | 9.3/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5723 字 阅读 →
论文解读

Elastic Time: Dynamic Frame Rate Bottlenecks for Neural Audio Coding

音频编解码 | 8.3/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4877 字 阅读 →
论文解读

FBK's Long-form SpeechLLMs for IWSLT 2026 Instruction Following

语音识别 | 6.5/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6105 字 阅读 →
论文解读

Generative AI and Copyright Infringement: A Legal-Technical Analysis of AI Music Generation Systems Under 17 U.S.C. Title 17

音乐生成 | 6.0/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4812 字 阅读 →
论文解读

Listening Like a Judge: A Music-Aware Framework for Automatic Singing Performance Evaluation

歌唱评估 | 8.8/10

 · 更新于 2026-09-07 · 约 16 分钟 · 7529 字 阅读 →
论文解读

Low Resource Multimodal Translation of Nepali Spoken Words into Emotion-Conditioned Sign Language Avatars

语音识别 | 5.3/10

 · 更新于 2026-09-07 · 约 23 分钟 · 11134 字 阅读 →
论文解读

Neural Speaker Diarization via Multilingual Training: Evaluation on Low-Resource Nepali-Hindi Speech

语音分离 | 5.5/10

 · 更新于 2026-09-07 · 约 15 分钟 · 7055 字 阅读 →
论文解读

Phonetic and semantic analyses of spoken corpora of Beijing and Taiwan Mandarin indicate that the neutral tone is a lexical tone

Phonetic and semantic analyses of spoken corpora of Beijing and Taiwan Mandarin indicate that the neutral tone is a lexical tone

 · 更新于 2026-09-07 · 约 1 分钟 · 59 字 阅读 →
论文解读

RedVox: Safety and Fairness Gaps in Speech Models Across Languages

基准测试 | 6.8/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4946 字 阅读 →
论文解读

SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages

语音识别 | 7.8/10

 · 更新于 2026-09-07 · 约 9 分钟 · 4420 字 阅读 →
论文解读

Soroll-IA: A Weakly Labeled Audio Dataset for Real-World Industrial Port Monitoring

音频事件检测 | 8.3/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5550 字 阅读 →
论文解读

Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents

知识蒸馏 | 6.7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5061 字 阅读 →