论文解读

SLM-TTA: A Framework for Test-Time Adaptation of Generative Spoken Language Models

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4556 字 阅读 →
论文解读

SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5250 字 阅读 →
论文解读

STACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio Codecs

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5581 字 阅读 →
论文解读

Streaming Speech Recognition with Decoder-Only Large Language Models and Latency Optimization

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4706 字 阅读 →
论文解读

Synthesized Data Selection via Score Distribution Matching for Te Reo Māori Automatic Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4219 字 阅读 →
论文解读

Synthetic Data Domain Adaptation for ASR via LLM-Based Text and Phonetic Respelling Augmentation

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4965 字 阅读 →
论文解读

TAGARELA - A Portuguese Speech Dataset from Podcasts

语音识别 语音合成 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3967 字 阅读 →
论文解读

Target-Speaker LLM-ASR with Speaker-Aware Speech Encoder

语音识别 | 8.8/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4444 字 阅读 →
论文解读

TASU: Text-only Alignment for Speech Understanding

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4344 字 阅读 →
论文解读

Teaching the Teachers: Boosting Unsupervised Domain Adaptation In Speech Recognition By Ensemble Update

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4428 字 阅读 →
论文解读

Three Seconds is Sufficient: A Multi-Pronged Framework for Model-Based Speaker Adaptation in ASR Under Data-Scarce Conditions

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4318 字 阅读 →
论文解读

TICL: Text-Embedding KNN for Speech in-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4710 字 阅读 →
论文解读

Tokenchain: A Discrete Speech Chain via Semantic Token Modeling

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6054 字 阅读 →
论文解读

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6217 字 阅读 →
论文解读

Towards Fair ASR for Second Language Speakers using Fairness Prompted Finetuning

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4138 字 阅读 →
论文解读

Towards Orthographically-Informed Evaluation of Speech Recognition Systems for Indian Languages

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3966 字 阅读 →
论文解读

Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER

语音识别 | 9.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4930 字 阅读 →
论文解读

Train Short, Infer Long: Speech-LLM Enables Zero-Shot Streamable Joint ASR and Diarization on Long Audio

说话人分离 | 9.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4925 字 阅读 →
论文解读

TTA: Transcribe, Translate and Alignment for Cross-Lingual Speech Representation

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5261 字 阅读 →
论文解读

UMA-SPLIT: Unimodal Aggregation for Both English and Mandarin Non-Autoregressive Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3808 字 阅读 →