论文解读

Lattice-Guided Consistency Regularization of Dual-Mode Transducers for Automatic Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3973 字 阅读 →
论文解读

Learning to Align with Unbalanced Optimal Transport in Linguistic Knowledge Transfer for ASR

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4273 字 阅读 →
论文解读

LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models Using in-the-wild Data

语音识别 语音翻译 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5574 字 阅读 →
论文解读

Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models

语音识别 | 6.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4691 字 阅读 →
论文解读

Leveraging Segment-Level Speech Representations for LLM-Based Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5225 字 阅读 →
论文解读

Leveraging Text-to-Speech and Voice Conversion as Data Augmentation for Alzheimer's Disease Detection from Spontaneous Speech

语音生物标志物 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5134 字 阅读 →
论文解读

Linguard: Authenticating Speech Recordings Using Speech Recognition and Watermark

音频安全 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4307 字 阅读 →
论文解读

Listen, But Don't Leak: Sensitive Data Protection for Privacy Aware Automatic Speech Recognition with Acoustic Triggers

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4187 字 阅读 →
论文解读

LLM-Based Post-ASR Error Correction for Disordered Speech

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5031 字 阅读 →
论文解读

LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech

基准测试 | 7.8/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3975 字 阅读 →
论文解读

LOTUSDIS: A Thai Far-Field Meeting Corpus for Robust Conversational ASR

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3735 字 阅读 →
论文解读

Medical ASR Enhancement by Domain-Specific Reinforcement Fine-Tuning

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4017 字 阅读 →
论文解读

Mind the Shift: Using Delta SSL Embeddings to Enhance Child ASR

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3909 字 阅读 →
论文解读

Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4158 字 阅读 →
论文解读

Mixture To Beamformed Mixture: Leveraging Beamformed Mixture As Weak-Supervision for Speech Enhancement and Noise-Robust ASR

语音增强 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4195 字 阅读 →
论文解读

Mixtures of Lightweight Articulatory Experts for Multilingual Asr

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3929 字 阅读 →
论文解读

MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3597 字 阅读 →
论文解读

Multilingual Supervised Pretraining with Lm-Assisted Decoding for Visual Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3573 字 阅读 →
论文解读

nGPT as a Scalable Architecture for Speech Recognition and Translation

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5927 字 阅读 →
论文解读

Noise-Robust AV-ASR Using Visual Features both in the Whisper Encoder and Decoder

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4647 字 阅读 →