论文解读

LOTUSDIS: A Thai Far-Field Meeting Corpus for Robust Conversational ASR

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3735 字 阅读 →
论文解读

Medical ASR Enhancement by Domain-Specific Reinforcement Fine-Tuning

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4017 字 阅读 →
论文解读

Mind the Shift: Using Delta SSL Embeddings to Enhance Child ASR

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3909 字 阅读 →
论文解读

Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4158 字 阅读 →
论文解读

Mixture To Beamformed Mixture: Leveraging Beamformed Mixture As Weak-Supervision for Speech Enhancement and Noise-Robust ASR

语音增强 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4195 字 阅读 →
论文解读

Mixtures of Lightweight Articulatory Experts for Multilingual Asr

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3929 字 阅读 →
论文解读

MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3597 字 阅读 →
论文解读

Multilingual Supervised Pretraining with Lm-Assisted Decoding for Visual Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3573 字 阅读 →
论文解读

nGPT as a Scalable Architecture for Speech Recognition and Translation

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5927 字 阅读 →
论文解读

Noise-Robust AV-ASR Using Visual Features both in the Whisper Encoder and Decoder

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4647 字 阅读 →
论文解读

OMNI-AVSR: Towards Unified Multimodal Speech Recognition With Large Language Models

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4649 字 阅读 →
论文解读

Online Register For Dual-Mode Self-Supervised Speech Models: Mitigating the Lack of Future Context

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5121 字 阅读 →
论文解读

PAC: Pronunciation-Aware Contextualized Large Language Model-Based Automatic Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4114 字 阅读 →
论文解读

Peeking Into the Future for Contextual Biasing

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5086 字 阅读 →
论文解读

PhoenixDSR: Phoneme-Guided and LLM-Enhanced Dysarthric Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5876 字 阅读 →
论文解读

Polynomial Mixing for Efficient Self-Supervised Speech Encoders

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4653 字 阅读 →
论文解读

Position-Invariant Fine-Tuning Of Speech Enhancement Models With Self-Supervised Speech Representations

语音增强 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4224 字 阅读 →
论文解读

Production-Scale Dynamic Vocabulary ASR Biasing with Word-Level FST and Robust Training

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3562 字 阅读 →
论文解读

Proficiency-Aware Adaptation and Data Augmentation for Robust L2 ASR

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3992 字 阅读 →
论文解读

Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4265 字 阅读 →