论文解读

Bayesian Low-Rank Factorization for Robust Model Adaptation

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4176 字 阅读 →
论文解读

BBPE16: UTF-16-Based Byte-Level Byte-Pair Encoding for Improved Multilingual Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3372 字 阅读 →
论文解读

BEST-RQ-based Self-Supervised Learning for Whisper Domain Adaptation

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5130 字 阅读 →
论文解读

BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4898 字 阅读 →
论文解读

Bridging the Front-End and Back-End for Robust ASR via Cross-Attention-Based U-Net

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4857 字 阅读 →
论文解读

CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5139 字 阅读 →
论文解读

Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs

基准测试 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4576 字 阅读 →
论文解读

CCST: Cross-Modal and Consistency-Aware Self-Training for Source-Free Unsupervised Domain Adaptation in Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6342 字 阅读 →
论文解读

Chunk-Wise Attention Transducers for Fast and Accurate Streaming Speech-to-Text

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4722 字 阅读 →
论文解读

Chunkwise Aligners for Streaming Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4405 字 阅读 →
论文解读

Confidence-Guided Error Correction for Disordered Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4863 字 阅读 →
论文解读

Content-Preserving Speech Representation Learning Via Adaptive Segment-Level Alignment

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5415 字 阅读 →
论文解读

Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5305 字 阅读 →
论文解读

Cross-Cultural Bias in Mel-Scale Representations: Evidence and Alternatives from Speech and Music

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4247 字 阅读 →
论文解读

Cross-Modal Bottleneck Fusion for Noise Robust Audio-Visual Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4237 字 阅读 →
论文解读

CTC-DID: CTC-Based Arabic Dialect Identification for Streaming Applications

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4764 字 阅读 →
论文解读

Decoder-Only Conformer with Modality-Aware Sparse Mixtures of Experts for ASR

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5265 字 阅读 →
论文解读

Do we really need self-attention for streaming automatic speech recognition?

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4519 字 阅读 →
论文解读

Domain-Aware Scheduling for ASR Fine-Tuning

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4493 字 阅读 →
论文解读

Emilia-NV: A Non-Verbal Speech Dataset with Word-Level Annotation for Human-Like Speech Modeling

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4512 字 阅读 →