论文解读

LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition

语音识别 | 9.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4128 字 阅读 →
论文解读

Qualitative Evaluation of Language Model Rescoring in Automatic Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 5 分钟 · 2252 字 阅读 →
论文解读

Selective Augmentation: Improving Universal Automatic Phonetic Transcription via G2P Bootstrapping

语音识别 | 6.0/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3160 字 阅读 →
论文解读

StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario

数据集 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3921 字 阅读 →
论文解读

Text-Utilization for Encoder-dominated Speech Recognition Models

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 6 分钟 · 2856 字 阅读 →
论文解读

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

语音对话系统 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3628 字 阅读 →
论文解读

A Personalized Real-Time Proactive Voice Memory Assistant

实时处理 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4384 字 阅读 →
论文解读

A Study of Data Selection Strategies for Pre-Training Self-Supervised Speech Models

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4254 字 阅读 →
论文解读

A Text-To-Text Alignment Algorithm for Better Evaluation of Modern Speech Recognition Systems

模型评估 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3720 字 阅读 →
论文解读

AccLID: Accent-aware Language Identification for Robust Multilingual Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4432 字 阅读 →
论文解读

Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5056 字 阅读 →
论文解读

Advanced modeling of interlanguage speech intelligibility benefit with L1-L2 multi-task learning using differentiable K-means for accent-robust discrete token-based ASR

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4144 字 阅读 →
论文解读

Advancing LLM-Based Multi-Channel Multi-Speaker Speech Recognition with Global Cross-Channel Attention and Sentence-Ordered First-In First-Out Serialized Output Training

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5058 字 阅读 →
论文解读

Advancing Semi-Supervised Child Speech Recognition with Omni-Temporal Classification under Label Noise

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4912 字 阅读 →
论文解读

Adversarial Fine-Tuning on Speech Foundation Model with Vulnerable Attention Consistency Regularization for Robust Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4509 字 阅读 →
论文解读

AISHELL6-Whisper: A Chinese Mandarin Audio-Visual Whisper Speech Dataset with Speech Recognition Baselines

语音识别 | 8.3/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4735 字 阅读 →
论文解读

An End-to-End Multimodal System for Subtitle Recognition and Chinese-Japanese Translation in Short Dramas

多模态模型 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5044 字 阅读 →
论文解读

Ara-BEST-RQ: Multi Dialectal Arabic SSL

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4587 字 阅读 →
论文解读

Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-text System

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5189 字 阅读 →
论文解读

Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4086 字 阅读 →