论文解读

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

语音识别 | 9.5/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5563 字 阅读 →
论文解读

Learning Robust Pair Confidence for Multimodal Emotion-Cause Pair Extraction

多模态模型 | 7.5/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6808 字 阅读 →
论文解读

Low-resource Language Discrimination Towards Chinese Dialects with Transfer learning and Data Augmentation

语音识别 | 5.5/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5615 字 阅读 →
论文解读

MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data

语音合成 | 7.9/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5712 字 阅读 →
论文解读

Mitigating Scoring Errors and Compensating for Nonverbal Subtests in Speech-Based Dementia Assessment

多模态模型 | 8/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6686 字 阅读 →
论文解读

Montreal Forced Aligner and the state of speech-to-text alignment in 2026

语音识别 | 7.5/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6323 字 阅读 →
论文解读

Native Active Perception as Reasoning for Omni-Modal Understanding

语音识别 | 9.1/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6616 字 阅读 →
论文解读

NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization

声源定位 | 7.8/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5653 字 阅读 →
论文解读

QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement

语音增强 | 7.1/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6928 字 阅读 →
论文解读

Reference-Based Recursive Least-Squares Mitigation of Real Interference in Stereo Audio Recordings

Reference-Based Recursive Least-Squares Mitigation of Real Interference in Stereo Audio Recordings

 · 更新于 2026-09-07 · 约 11 分钟 · 5236 字 阅读 →
论文解读

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

语音合成 | 7.3/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5547 字 阅读 →
论文解读

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs

语音合成 | 7.4/10

 · 更新于 2026-09-07 · 约 15 分钟 · 7228 字 阅读 →
论文解读

Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings

语音识别 | 6.5/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6503 字 阅读 →
论文解读

Risk Stratification for ICU Delirium using Pervasive Ambient Sensing Information

多模态模型 | 6.5/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4563 字 阅读 →
论文解读

Scoring Backends Matter More Than Pooling: A Systematic Study of Training-Free Anomalous Sound Detection under Domain Shift

Scoring Backends Matter More Than Pooling: A Systematic Study of Training-Free Anomalous Sound Detection under Domain Shift

 · 更新于 2026-09-07 · 约 10 分钟 · 4814 字 阅读 →
论文解读

SingFox: A Multi-Lingual Singfake Detection Corpus

语音伪造检测 | 5.4/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6765 字 阅读 →
论文解读

Speech-Driven End-to-End Language Discrimination towards Chinese Dialects

语音识别 | 5.8/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6225 字 阅读 →
论文解读

ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection

强化学习 | 6.3/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4793 字 阅读 →
论文解读

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs

多模态模型 | 8.8/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5845 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-18

共分析 36 篇语音/AI 论文

 · 更新于 2026-09-07 · 约 105 分钟 · 52107 字 阅读 →