论文解读

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

语音识别 | 8.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6046 字 阅读 →
论文解读

Bridging Traditional Explainability Methods and Multimodal Multilingual Models: An XAI-Based Analysis

语音识别 | 5.4/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8872 字 阅读 →
论文解读

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7716 字 阅读 →
论文解读

Factors affecting ASR performance: A study using state of the art ASR models in Indic Languages

语音识别 | 6.9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5995 字 阅读 →
论文解读

Is Text All You Need? Text as a Universal Information Bottleneck for Speech LLMs

语音识别 | 7.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5311 字 阅读 →
论文解读

Overcoming Decoder Inconsistencies in Whisper for Dravidian and Low-Resource Languages

语音识别 | 6.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6234 字 阅读 →
论文解读

Parameter-Efficient Continual Learning for Automatic Speech Recognition

语音识别 | 8.1/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5336 字 阅读 →
论文解读

Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER)

语音识别 | 8.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6640 字 阅读 →
论文解读

Assessing True Generalisability of Audio-Visual Speech Recognisers

语音识别 | 9.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6841 字 阅读 →
论文解读

Contrastive Training with LLM-generated Near-Misses for Robust Code-Switching Speech Recognition

语音识别 | 7.1/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5503 字 阅读 →
论文解读

Hearing the Unspoken: Language Model Priors for Acoustic Adversarial Attacks

语音识别 | 9.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5472 字 阅读 →
论文解读

Phonetic Error Analysis of Raw Waveform Acoustic Models

语音识别 | 7.6/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4614 字 阅读 →
论文解读

Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders

语音识别 | 7.9/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5309 字 阅读 →
论文解读

Age-Aware Adapter Tuning for Children's Speech Recognition

语音识别 | 8.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5531 字 阅读 →
论文解读

Automatic Labelling of Speech Translation Errors

语音识别 | 6.1/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4750 字 阅读 →
论文解读

Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5073 字 阅读 →
论文解读

Beyond WER: A Paired Acoustic Stress Test for Ambient Clinical Scribes

语音识别 | 7.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5672 字 阅读 →
论文解读

CoSTA: Cognitive-State-Conditioned TTS Data Augmentation Using ASR Transcripts for Alzheimer's Disease Detection

语音合成 | 6.5/10

 · 更新于 2026-09-25 · 约 4 分钟 · 1509 字 阅读 →
论文解读

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation

语音合成 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6113 字 阅读 →
论文解读

FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7310 字 阅读 →