论文解读

A study on weakly-supervised training approaches for phoneme-level pronunciation scoring

语音识别 | 9.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6438 字 阅读 →
论文解读

Articulatory strategy as a source of variation in acoustic vowel dynamics

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5549 字 阅读 →
论文解读

Convex Low-resource Accent-Robust Language Detection in Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6301 字 阅读 →
论文解读

StepAudio 2.5 Technical Report

统一音频模型 | 8.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6187 字 阅读 →
论文解读

Word-Level Modeling with Alignment-Aware Acoustic Fusion for Text-Assisted Intelligibility Prediction in Listeners with Hearing Loss

语音质量评估 | 7.7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6639 字 阅读 →
论文解读

Convex Low-resource Accent-Robust Language Detection in Speech Recognition

** | 7.5/10

 · 更新于 2026-09-25 · 约 1 分钟 · 78 字 阅读 →
论文解读

Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German

语音识别 | 6.8/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7727 字 阅读 →
论文解读

FormalASR: End-to-End Spoken Chinese to Formal Text

语音识别 | 8.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7790 字 阅读 →
论文解读

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

语音识别 | 9.3/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7818 字 阅读 →
论文解读

SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR

语音识别 | 8.3/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6832 字 阅读 →
论文解读

Can Large Language Models Reliably Correct Errors in Low-Resource ASR? A Contamination-Aware Case Study on West Frisian

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9491 字 阅读 →
论文解读

FormalASR: End-to-End Spoken Chinese to Formal Text

语音识别 | 6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7432 字 阅读 →
论文解读

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

语音识别 | 6.8/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10575 字 阅读 →
论文解读

Contextual Biasing for Streaming ASR via CTC-based Word Spotting

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7082 字 阅读 →
论文解读

MedASR: An Open-Source Model for High-Accuracy Medical Dictation

语音识别 | 7.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8271 字 阅读 →
论文解读

Sometin Beta Pass Notin (SBPN): Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation

语音识别 | 6.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7571 字 阅读 →
论文解读

UrduSpeech: A 156-Hour Urdu Speech Corpus with 12-Dimension Paralinguistic Annotations

语音识别 | 7/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7835 字 阅读 →
论文解读

Improving Automatic Speech Recognition for Speakers Treated for Oral Cancer using Data Augmentation and LLM Error Correction

语音识别 | 6/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8703 字 阅读 →
论文解读

Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization

语音识别 说话人分离 | 7.2/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8034 字 阅读 →
论文解读

A Calculus-Based Framework for Determining Vocabulary Size in End-to-End ASR

语音识别 | 3.9/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8601 字 阅读 →