论文解读

Hearing the Unspoken: Language Model Priors for Acoustic Adversarial Attacks

语音识别 | 9.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5472 字 阅读 →
论文解读

Phonetic Error Analysis of Raw Waveform Acoustic Models

语音识别 | 7.6/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4614 字 阅读 →
论文解读

Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders

语音识别 | 7.9/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5309 字 阅读 →
论文解读

Age-Aware Adapter Tuning for Children's Speech Recognition

语音识别 | 8.4/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5531 字 阅读 →
论文解读

Automatic Labelling of Speech Translation Errors

语音识别 | 6.1/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4750 字 阅读 →
论文解读

Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5073 字 阅读 →
论文解读

Beyond WER: A Paired Acoustic Stress Test for Ambient Clinical Scribes

语音识别 | 7.1/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5672 字 阅读 →
论文解读

CoSTA: Cognitive-State-Conditioned TTS Data Augmentation Using ASR Transcripts for Alzheimer's Disease Detection

语音合成 | 6.5/10

 · 更新于 2026-09-06 · 约 4 分钟 · 1509 字 阅读 →
论文解读

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation

语音合成 | 7.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6113 字 阅读 →
论文解读

FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7310 字 阅读 →
论文解读

M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition

语音识别 | 9/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6425 字 阅读 →
论文解读

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

语音识别 | 8.4/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6696 字 阅读 →
论文解读

Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition

语音识别 | 6.9/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4724 字 阅读 →
论文解读

Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios

语音翻译 | 8.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5825 字 阅读 →
论文解读

Revisiting Lexicon Evaluation in Unsupervised Word Discovery

语音识别 | 1/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5346 字 阅读 →
论文解读

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

语音识别 | 7.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6366 字 阅读 →
论文解读

Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs

语音识别 | 5.9/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5559 字 阅读 →
论文解读

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

语音编码 | 8.8/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7045 字 阅读 →
论文解读

Multilingual Long-Form Speech Instruction Following: KIT's Submission to IWSLT 2026

语音识别 | 10/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6280 字 阅读 →
论文解读

Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy

语音识别 | 8.6/10

 · 更新于 2026-09-06 · 约 3 分钟 · 1214 字 阅读 →