论文解读

VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track

音频问答 | 3.9/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5843 字 阅读 →
论文解读

VoxCPM2 Technical Report

语音合成 | 9.5/10

 · 更新于 2026-09-09 · 约 15 分钟 · 7125 字 阅读 →
论文解读

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

多模态模型 | 6.4/10

 · 更新于 2026-09-09 · 约 8 分钟 · 3789 字 阅读 →
论文解读

Where Rectified Flows Leak: Characterising Membership Signals Along the Interpolation Path

音频生成 | 8.7/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6031 字 阅读 →
论文解读

Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders

语音识别 | 7.9/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5309 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-08

共分析 38 篇语音/AI 论文

 · 更新于 2026-09-09 · 约 108 分钟 · 53770 字 阅读 →
论文解读

A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing

A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing

 · 更新于 2026-09-09 · 约 13 分钟 · 6025 字 阅读 →
论文解读

Age-Aware Adapter Tuning for Children's Speech Recognition

语音识别 | 8.4/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5531 字 阅读 →
论文解读

An ERP Study on Recursive Locative Processing in Mandarin-Speaking Children with Autism

An ERP Study on Recursive Locative Processing in Mandarin-Speaking Children with Autism

 · 更新于 2026-09-09 · 约 12 分钟 · 5787 字 阅读 →
论文解读

An Ultra-Low-Bitrate Neural Speech Codec with Plain-to-Pseudo Synergistic Vector Quantization

语音合成 | 7.7/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5324 字 阅读 →
论文解读

Audio Interaction Model

流式处理 | 9.8/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5688 字 阅读 →
论文解读

Automatic Labelling of Speech Translation Errors

语音识别 | 6.1/10

 · 更新于 2026-09-09 · 约 10 分钟 · 4750 字 阅读 →
论文解读

Beyond Generative Decoding: Discriminative Hidden-State Readout from a Native Omni-Modal LLM for Multimodal Sentiment Analysis

多模态模型 | 5.3/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6789 字 阅读 →
论文解读

Beyond Text Following: Repairable Arbitration Reversals in Audio-Language Models

音频问答 | 6.4/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6582 字 阅读 →
论文解读

Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-09 · 约 11 分钟 · 5073 字 阅读 →
论文解读

Beyond WER: A Paired Acoustic Stress Test for Ambient Clinical Scribes

语音识别 | 7.1/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5672 字 阅读 →
论文解读

CoSTA: Cognitive-State-Conditioned TTS Data Augmentation Using ASR Transcripts for Alzheimer's Disease Detection

语音合成 | 6.5/10

 · 更新于 2026-09-09 · 约 4 分钟 · 1509 字 阅读 →
论文解读

DBHN-Net: Dual-Branch Hybrid Neural Network For Low-Complexity Monaural Speech Enhancement

语音增强 | 5.4/10

 · 更新于 2026-09-09 · 约 14 分钟 · 6912 字 阅读 →
论文解读

Do speech foundation models perceive speaker similarity as humans do?

说话人识别 | 6.3/10

 · 更新于 2026-09-09 · 约 13 分钟 · 6169 字 阅读 →
论文解读

Domain-Aware Mispronunciation Detection and Diagnosis Using Language-Specific Statistical Graphs

图神经网络 | 6.6/10

 · 更新于 2026-09-09 · 约 12 分钟 · 5824 字 阅读 →