论文解读

Can Large Language Models Reliably Correct Errors in Low-Resource ASR? A Contamination-Aware Case Study on West Frisian

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9491 字 阅读 →
论文解读

FormalASR: End-to-End Spoken Chinese to Formal Text

语音识别 | 6/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7432 字 阅读 →
论文解读

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

语音识别 | 6.8/10

 · 更新于 2026-09-06 · 约 22 分钟 · 10575 字 阅读 →
论文解读

Contextual Biasing for Streaming ASR via CTC-based Word Spotting

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7082 字 阅读 →
论文解读

MedASR: An Open-Source Model for High-Accuracy Medical Dictation

语音识别 | 7.9/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8271 字 阅读 →
论文解读

Sometin Beta Pass Notin (SBPN): Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation

语音识别 | 6.2/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7571 字 阅读 →
论文解读

UrduSpeech: A 156-Hour Urdu Speech Corpus with 12-Dimension Paralinguistic Annotations

语音识别 | 7/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7835 字 阅读 →
论文解读

Improving Automatic Speech Recognition for Speakers Treated for Oral Cancer using Data Augmentation and LLM Error Correction

语音识别 | 6/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8703 字 阅读 →
论文解读

Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization

语音识别 说话人分离 | 7.2/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8034 字 阅读 →
论文解读

A Calculus-Based Framework for Determining Vocabulary Size in End-to-End ASR

语音识别 | 3.9/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8601 字 阅读 →
论文解读

Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8162 字 阅读 →
论文解读

Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8776 字 阅读 →
论文解读

WARDEN: Endangered Indigenous Language Transcription and Translation with 6 Hours of Training Data

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7464 字 阅读 →
论文解读

Chunkwise Aligners for Streaming Speech Recognition

语音识别 | 6.3/10

 · 更新于 2026-09-06 · 约 20 分钟 · 9896 字 阅读 →
论文解读

Mechanistic Interpretability of ASR models using Sparse Autoencoders

语音识别 | 5.5/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8102 字 阅读 →
论文解读

Too Good to Be True: A Study on Modern Automatic Speech Recognition for the Evaluation of Speech Enhancement

语音增强 | 6.6/10

 · 更新于 2026-09-06 · 约 21 分钟 · 10496 字 阅读 →
论文解读

Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization

语音识别 说话人日志 | 5.5/10

 · 更新于 2026-09-06 · 约 21 分钟 · 10162 字 阅读 →
论文解读

Dolphin-CN-Dialect: Where Chinese Dialects Matter

语音识别 | 5.5/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8629 字 阅读 →
论文解读

Responsible Benchmarking of Fairness for Automatic Speech Recognition

语音识别 | 5.0/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7970 字 阅读 →
论文解读

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models

语音识别 | 6.0/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9094 字 阅读 →