论文解读

An Analysis of the Effectiveness of Synthetic Speech Data for ASR Fine-tuning in Selected Indic Languages

语音识别 | 5.7/10

 · 更新于 2026-09-06 · 约 25 分钟 · 12144 字 阅读 →
论文解读

Are you speaking my languages? On spoken language adherence in multimodal LLMs

语音识别 | 8/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4530 字 阅读 →
论文解读

Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning

语音识别 | 8.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5868 字 阅读 →
论文解读

Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4484 字 阅读 →
论文解读

MLLP-VRAIN UPV system for the IWSLT 2026 Simultaneous Speech Translation task

语音识别 | 6.9/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7669 字 阅读 →
论文解读

Next-Turn: Duration-Aware Streaming Endpoint Detection via Time-to-Next-Speech-Onset Prediction

语音合成 | 7.9/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5858 字 阅读 →
论文解读

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI

语音识别 | 7.6/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6566 字 阅读 →
论文解读

When Multiple Scripts Matter: Evaluating ASR in Clinical Settings

语音识别 | 9.1/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6513 字 阅读 →
论文解读

AP-GRPO: Anchor-Gated Phonetic Alignment with Policy Optimization for Pathological Speech Reconstruction

语音识别 | 7.4/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6856 字 阅读 →
论文解读

ArtBoost: Synthetic Articulatory Data Augmentation for Acoustic-to-Articulatory Inversion

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5884 字 阅读 →
论文解读

ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition

语音识别 | 8.3/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4766 字 阅读 →
论文解读

Bridging the SEA Gap: An Initial Benchmark for Neural Audio Codec-Synthesized Speech Deepfakes in South-East Asian Languages

语音合成 | 8.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5381 字 阅读 →
论文解读

Confidence Score Guided Incremental and Speaker Adaptive Pseudo-Labeling for Semi-Supervised Elderly Speech Recognition

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6027 字 阅读 →
论文解读

Decoding while Adapting: Zero-Shot Online Speaker Adaptation via Audio-Textual Prompts for Elderly Speech Recognition

语音识别 | 6.8/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6816 字 阅读 →
论文解读

EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning

音频问答 | 6.1/10

 · 更新于 2026-09-06 · 约 22 分钟 · 10753 字 阅读 →
论文解读

From Awareness to Adherence: Bridging the Context Gap in Spoken Dialogue Systems via Context-Aware Decoding

语音识别 | 6.7/10

 · 更新于 2026-09-06 · 约 5 分钟 · 2144 字 阅读 →
论文解读

MambAdapter: Lightweight Mamba-Based Adapters for Parameter-Efficient Transfer Learning in Speech and Audio

语音识别 | 8.9/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4872 字 阅读 →
论文解读

ROMPAR: Morphological Completion and Demographic Unlearning for Romanian-Accented Speech Recognition

语音识别 | 6.2/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4459 字 阅读 →
论文解读

Scaling Human and G2P Supervision for Robust Phonetic Transcription

语音识别 | 7.6/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4873 字 阅读 →
论文解读

BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM

语音合成 | 9/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5563 字 阅读 →