论文解读

Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5056 字 阅读 →
论文解读

Advanced modeling of interlanguage speech intelligibility benefit with L1-L2 multi-task learning using differentiable K-means for accent-robust discrete token-based ASR

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4144 字 阅读 →
论文解读

Advancing LLM-Based Multi-Channel Multi-Speaker Speech Recognition with Global Cross-Channel Attention and Sentence-Ordered First-In First-Out Serialized Output Training

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5058 字 阅读 →
论文解读

Advancing Semi-Supervised Child Speech Recognition with Omni-Temporal Classification under Label Noise

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4912 字 阅读 →
论文解读

Adversarial Fine-Tuning on Speech Foundation Model with Vulnerable Attention Consistency Regularization for Robust Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4509 字 阅读 →
论文解读

AISHELL6-Whisper: A Chinese Mandarin Audio-Visual Whisper Speech Dataset with Speech Recognition Baselines

语音识别 | 8.3/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4735 字 阅读 →
论文解读

An End-to-End Multimodal System for Subtitle Recognition and Chinese-Japanese Translation in Short Dramas

多模态模型 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5044 字 阅读 →
论文解读

Ara-BEST-RQ: Multi Dialectal Arabic SSL

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4587 字 阅读 →
论文解读

Attention2Probability: Attention-Driven Terminology Probability Estimation for Robust Speech-to-text System

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5189 字 阅读 →
论文解读

Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4086 字 阅读 →
论文解读

Bayesian Low-Rank Factorization for Robust Model Adaptation

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4176 字 阅读 →
论文解读

BBPE16: UTF-16-Based Byte-Level Byte-Pair Encoding for Improved Multilingual Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3372 字 阅读 →
论文解读

BEST-RQ-based Self-Supervised Learning for Whisper Domain Adaptation

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5130 字 阅读 →
论文解读

BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4898 字 阅读 →
论文解读

Bridging the Front-End and Back-End for Robust ASR via Cross-Attention-Based U-Net

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4857 字 阅读 →
论文解读

CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5139 字 阅读 →
论文解读

Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs

基准测试 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4576 字 阅读 →
论文解读

CCST: Cross-Modal and Consistency-Aware Self-Training for Source-Free Unsupervised Domain Adaptation in Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6342 字 阅读 →
论文解读

Chunk-Wise Attention Transducers for Fast and Accurate Streaming Speech-to-Text

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4722 字 阅读 →
论文解读

Chunkwise Aligners for Streaming Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4405 字 阅读 →