论文解读

CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS

语音编辑 | 7.2/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7769 字 阅读 →
论文解读

FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions

语音识别 | 5.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5287 字 阅读 →
论文解读

PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech

语音合成 | 6.5/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8339 字 阅读 →
论文解读

Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems

语音识别 | 6.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6086 字 阅读 →
论文解读

Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5492 字 阅读 →
论文解读

Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization

语音识别 | 6.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6387 字 阅读 →
论文解读

Why Can't They Remember? Uncovering Representation and Retrieval Bottlenecks in Multi-Turn Acoustic Memory

语音识别 | 7/10

 · 更新于 2026-09-06 · 约 4 分钟 · 1800 字 阅读 →
论文解读

Multilingual Phonological Feature Recognition with Self-Supervised Speech Models

语音识别 | 7.7/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5589 字 阅读 →
论文解读

Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems

语音识别 | 9.6/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8124 字 阅读 →
论文解读

Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems

语音识别 | 6.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4440 字 阅读 →
论文解读

A study on weakly-supervised training approaches for phoneme-level pronunciation scoring

语音识别 | 9.7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6438 字 阅读 →
论文解读

Articulatory strategy as a source of variation in acoustic vowel dynamics

语音识别 | 8.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5549 字 阅读 →
论文解读

Convex Low-resource Accent-Robust Language Detection in Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6301 字 阅读 →
论文解读

StepAudio 2.5 Technical Report

统一音频模型 | 8.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6187 字 阅读 →
论文解读

Word-Level Modeling with Alignment-Aware Acoustic Fusion for Text-Assisted Intelligibility Prediction in Listeners with Hearing Loss

语音质量评估 | 7.7/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6639 字 阅读 →
论文解读

Convex Low-resource Accent-Robust Language Detection in Speech Recognition

** | 7.5/10

 · 更新于 2026-09-06 · 约 1 分钟 · 78 字 阅读 →
论文解读

Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German

语音识别 | 6.8/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7727 字 阅读 →
论文解读

FormalASR: End-to-End Spoken Chinese to Formal Text

语音识别 | 8.2/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7790 字 阅读 →
论文解读

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

语音识别 | 9.3/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7818 字 阅读 →
论文解读

SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR

语音识别 | 8.3/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6832 字 阅读 →