论文解读

Using Phonological-Level Wav2Vec2 for Mandarin Automatic Mispronunciation Detection and Diagnosis

语音识别 | 8.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6199 字 阅读 →
论文解读

Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi

语音识别 | 6.7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5011 字 阅读 →
论文解读

Word Lengthening as a Function of Utterance Position: A Multi-Corpus Study

语音合成 | 8.1/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4828 字 阅读 →
论文解读

A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5716 字 阅读 →
论文解读

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine

语音合成 | 7.4/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5736 字 阅读 →
论文解读

Analyzing Language and Geographical Variation in Speech Representations Across 60 Indic Languages

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5299 字 阅读 →
论文解读

Cross-Dataset, Age, and Gender Generalization: A Comprehensive Analysis of Fine-Tuning Strategies for Low-Resource Children's ASR

语音识别 | 6.7/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5701 字 阅读 →
论文解读

Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech

语音识别 | 7.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5725 字 阅读 →
论文解读

Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 23 分钟 · 11077 字 阅读 →
论文解读

Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning

语音识别 | 8.7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5408 字 阅读 →
论文解读

ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion

语音合成 | 6.2/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7884 字 阅读 →
论文解读

S-JEPA : Soft Clustering Anchors for Self-Supervised Speech Representation Learning

语音识别 | 8.7/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5785 字 阅读 →
论文解读

Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models

语音识别 | 8.3/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4610 字 阅读 →
论文解读

DASH: Dual-View Self-Distillation with Multi-Layer Hidden Representations for Robust Speech Recognition

语音识别 | 6.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5965 字 阅读 →
论文解读

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

语音识别 | 9.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5563 字 阅读 →
论文解读

Low-resource Language Discrimination Towards Chinese Dialects with Transfer learning and Data Augmentation

语音识别 | 5.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5615 字 阅读 →
论文解读

Montreal Forced Aligner and the state of speech-to-text alignment in 2026

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6323 字 阅读 →
论文解读

Native Active Perception as Reasoning for Omni-Modal Understanding

语音识别 | 9.1/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6616 字 阅读 →
论文解读

Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6503 字 阅读 →
论文解读

Speech-Driven End-to-End Language Discrimination towards Chinese Dialects

语音识别 | 5.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6225 字 阅读 →