论文解读

Enhancing Law-Enforcement Audio Transcription: A LoRA-Based Adaptation of Whisper for BWC Footage

语音识别 | 4.8/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5460 字 阅读 →
论文解读

Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens

语音识别 | 6.3/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8991 字 阅读 →
论文解读

Voice Memory for Agentic Speech Recognition

语音识别 | 8.2/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8134 字 阅读 →
论文解读

Evaluation of forced alignment of code-mixed speech: the case of Hindi-English

语音识别 | 5.4/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4432 字 阅读 →
论文解读

SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies

语音识别 | 7.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5520 字 阅读 →
论文解读

Towards Operational Conversational Intelligence: A Speech Intelligence Framework

说话人日志 | 6.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6520 字 阅读 →
论文解读

Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance

语音识别 | 7.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5808 字 阅读 →
论文解读

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

说话人日志 | 7.1/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8448 字 阅读 →
论文解读

MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

语音识别 | 6.7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4906 字 阅读 →
论文解读

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding

语音识别 | 6.6/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9351 字 阅读 →
论文解读

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond

语音识别 | 8.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6955 字 阅读 →
论文解读

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

语音识别 | 8.1/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7224 字 阅读 →
论文解读

From Read Speech to Spoken Digits: A Task-Specific Evaluation of Speech Privacy With Informed Attackers

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9074 字 阅读 →
论文解读

Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin

语音识别 | 6.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7221 字 阅读 →
论文解读

VibeVoice-ASR-BitNet Technical Report

语音识别 | 7.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6193 字 阅读 →
论文解读

Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8304 字 阅读 →
论文解读

Constrained CTC Decoding for Efficient Diacritic Restoration

语音识别 | 7.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6374 字 阅读 →
论文解读

From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4601 字 阅读 →
论文解读

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

语音识别 | 7.1/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8698 字 阅读 →
论文解读

What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

语音识别 | 4.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6414 字 阅读 →