论文解读

Three Seconds is Sufficient: A Multi-Pronged Framework for Model-Based Speaker Adaptation in ASR Under Data-Scarce Conditions

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4318 字 阅读 →
论文解读

TICL: Text-Embedding KNN for Speech in-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4710 字 阅读 →
论文解读

Tokenchain: A Discrete Speech Chain via Semantic Token Modeling

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6054 字 阅读 →
论文解读

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6217 字 阅读 →
论文解读

Towards Fair ASR for Second Language Speakers using Fairness Prompted Finetuning

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4138 字 阅读 →
论文解读

Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3966 字 阅读 →
论文解读

Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER

语音识别 | 9.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4930 字 阅读 →
论文解读

Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio

说话人分离 | 9.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4925 字 阅读 →
论文解读

TTA: Transcribe, Translate and Alignment for Cross-Lingual Speech Representation

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5261 字 阅读 →
论文解读

UMA-SPLIT: Unimodal Aggregation for Both English and Mandarin Non-Autoregressive Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3808 字 阅读 →
论文解读

Variational Low-Rank Adaptation for Personalized Impaired Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5297 字 阅读 →
论文解读

Voting-Based Pitch Estimation with Temporal and Frequential Alignment and Correlation Aware Selection

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5178 字 阅读 →
论文解读

WAV2LEV: Predicting Levenshtein Edit Operation Sequences For Fine-Grained Estimation of Automatic Speech Recognition Error

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4224 字 阅读 →
论文解读

Whisper-FEST: Single-Channel Far-Field Enhanced Speech-to-text without Parallel Data

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4962 字 阅读 →
论文解读

Whisper-MLA: Reducing GPU Memory Consumption of ASR Models Based on MHA2MLA Conversion

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4484 字 阅读 →
论文解读

Whisper: Courtside Edition - Enhancing ASR Performance through LLM-Driven Context Generation

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3477 字 阅读 →
论文解读

WhisperPipe: A Resource-Efficient Streaming Architecture for Real-Time Automatic Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3139 字 阅读 →
论文解读

Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-Resource Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3795 字 阅读 →
论文解读

Z-Scores: A Metric for Linguistically Assessing Disfluency Removal

模型评估 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4710 字 阅读 →
论文解读

RAS: a Reliability Oriented Metric for Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4410 字 阅读 →