论文解读

RAS: a Reliability Oriented Metric for Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4136 字 阅读 →
论文解读

Reducing Prompt Sensitivity in LLM-Based Speech Recognition Through Learnable Projection

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4818 字 阅读 →
论文解读

Reference Microphone Selection for Guided Source Separation Based on The Normalized L-P Norm

语音增强 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3975 字 阅读 →
论文解读

Relative Time Intervals Representation For Word-Level Timestamping With Masked Training

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4823 字 阅读 →
论文解读

RLBR: Reinforcement Learning with Biasing Rewards for Contextual Speech Large Language Models

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3657 字 阅读 →
论文解读

Robust Accent Identification via Voice Conversion and Non-Timbral Embeddings

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3282 字 阅读 →
论文解读

Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4626 字 阅读 →
论文解读

SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5790 字 阅读 →
论文解读

SED: Structural Entropy Based Speech Discretization for Discrete Token-Based ASR

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4993 字 阅读 →
论文解读

Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3356 字 阅读 →
论文解读

SLM-TTA: A Framework for Test-Time Adaptation of Generative Spoken Language Models

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4556 字 阅读 →
论文解读

SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5250 字 阅读 →
论文解读

STACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio Codecs

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5581 字 阅读 →
论文解读

Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4706 字 阅读 →
论文解读

Synthesized Data Selection via Score Distribution Matching for Te Reo Māori Automatic Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4219 字 阅读 →
论文解读

Synthetic Data Domain Adaptation for ASR via LLM-Based Text and Phonetic Respelling Augmentation

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4965 字 阅读 →
论文解读

TAGARELA - A Portuguese Speech Dataset from Podcasts

语音识别 语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3967 字 阅读 →
论文解读

Target-Speaker LLM-ASR with Speaker-Aware Speech Encoder

语音识别 | 8.8/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4444 字 阅读 →
论文解读

TASU: Text-only Alignment for Speech Understanding

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4344 字 阅读 →
论文解读

Teaching the Teachers: Boosting Unsupervised Domain Adaptation In Speech Recognition By Ensemble Update

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4428 字 阅读 →