论文解读

ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion

语音合成 | 6.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7884 字 阅读 →
论文解读

S-JEPA : Soft Clustering Anchors for Self-Supervised Speech Representation Learning

语音识别 | 8.7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5785 字 阅读 →
论文解读

Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models

语音识别 | 8.3/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4610 字 阅读 →
论文解读

DASH: Dual-View Self-Distillation with Multi-Layer Hidden Representations for Robust Speech Recognition

语音识别 | 6.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5965 字 阅读 →
论文解读

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

语音识别 | 9.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5563 字 阅读 →
论文解读

Low-resource Language Discrimination Towards Chinese Dialects with Transfer learning and Data Augmentation

语音识别 | 5.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5615 字 阅读 →
论文解读

Montreal Forced Aligner and the state of speech-to-text alignment in 2026

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6323 字 阅读 →
论文解读

Native Active Perception as Reasoning for Omni-Modal Understanding

语音识别 | 9.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6616 字 阅读 →
论文解读

Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6503 字 阅读 →
论文解读

Speech-Driven End-to-End Language Discrimination towards Chinese Dialects

语音识别 | 5.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6225 字 阅读 →
论文解读

An Analysis of the Effectiveness of Synthetic Speech Data for ASR Fine-tuning in Selected Indic Languages

语音识别 | 5.7/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12144 字 阅读 →
论文解读

Are you speaking my languages? On spoken language adherence in multimodal LLMs

语音识别 | 8/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4530 字 阅读 →
论文解读

Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5868 字 阅读 →
论文解读

Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4484 字 阅读 →
论文解读

MLLP-VRAIN UPV system for the IWSLT 2026 Simultaneous Speech Translation task

语音识别 | 6.9/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7669 字 阅读 →
论文解读

Next-Turn: Duration-Aware Streaming Endpoint Detection via Time-to-Next-Speech-Onset Prediction

语音合成 | 7.9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5858 字 阅读 →
论文解读

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI

语音识别 | 7.6/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6566 字 阅读 →
论文解读

When Multiple Scripts Matter: Evaluating ASR in Clinical Settings

语音识别 | 9.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6513 字 阅读 →
论文解读

AP-GRPO: Anchor-Gated Phonetic Alignment with Policy Optimization for Pathological Speech Reconstruction

语音识别 | 7.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6856 字 阅读 →
论文解读

ArtBoost: Synthetic Articulatory Data Augmentation for Acoustic-to-Articulatory Inversion

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5884 字 阅读 →