论文解读

Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR

语音识别 | 8.3/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5961 字 阅读 →
论文解读

Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models

语音识别 | 9.6/10

 · 更新于 2026-09-06 · 约 4 分钟 · 1866 字 阅读 →
论文解读

MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7026 字 阅读 →
论文解读

The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models

语音识别 | 8.2/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5775 字 阅读 →
论文解读

Unsupervised Approaches for Global Prosodic Embedding Extraction

语音合成 | 7.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6086 字 阅读 →
论文解读

Balancing ASR and diarization in end-to-end LLMs for multi-talker speech recognition

语音识别 | 7.1/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5831 字 阅读 →
论文解读

Positional Encoding in the Context of Memristor-Based Analog Computation for Automatic Speech Recognition

语音识别 | 8/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4178 字 阅读 →
论文解读

PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue

语音合成 | 8.1/10

 · 更新于 2026-09-06 · 约 22 分钟 · 10641 字 阅读 →
论文解读

Evaluating Bias in Phoneme-Based Automatic Speech Recognition Systems: An Analysis of IPA Transcription Models

语音识别 | 8.8/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5179 字 阅读 →
论文解读

Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains

语音识别 | 9.1/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4665 字 阅读 →
论文解读

I Understand How You Feel: Enhancing Deeper Emotional Support Through Multilingual Emotional Validation in Dialogue System

语音识别 | 5.8/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7227 字 阅读 →
论文解读

Massive Open-Vocabulary Keyword Spotting

语音识别 | 9.8/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5180 字 阅读 →
论文解读

Pretrained self-supervised speech models can recognize unseen consonants

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4929 字 阅读 →
论文解读

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

语音合成 | 7.9/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4632 字 阅读 →
论文解读

Towards Data-free and Training-free Compression for Speech Foundation Models Using Parameter Clustering

语音识别 | 6.4/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5326 字 阅读 →
论文解读

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6593 字 阅读 →
论文解读

AuRA: Internalizing Audio Understanding into LLMs as LoRA

语音问答 | 7.5/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3100 字 阅读 →
论文解读

Entropy-Aware Domain-Routed Mixture-of-Experts Speech-LLM Framework: A Case Study of Multi-Domain Child-Adult ASR

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6225 字 阅读 →
论文解读

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6823 字 阅读 →
论文解读

GC-LoRA: Gated Convolutional LoRA for Parameter-Efficient Acoustic Adaptation

语音识别 | 7.6/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4079 字 阅读 →