论文解读

ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition

语音识别 | 8.3/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4766 字 阅读 →
论文解读

Bridging the SEA Gap: An Initial Benchmark for Neural Audio Codec-Synthesized Speech Deepfakes in South-East Asian Languages

语音合成 | 8.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5381 字 阅读 →
论文解读

Confidence Score Guided Incremental and Speaker Adaptive Pseudo-Labeling for Semi-Supervised Elderly Speech Recognition

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6027 字 阅读 →
论文解读

Decoding while Adapting: Zero-Shot Online Speaker Adaptation via Audio-Textual Prompts for Elderly Speech Recognition

语音识别 | 6.8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6816 字 阅读 →
论文解读

EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning

音频问答 | 6.1/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10753 字 阅读 →
论文解读

From Awareness to Adherence: Bridging the Context Gap in Spoken Dialogue Systems via Context-Aware Decoding

语音识别 | 6.7/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2144 字 阅读 →
论文解读

MambAdapter: Lightweight Mamba-Based Adapters for Parameter-Efficient Transfer Learning in Speech and Audio

语音识别 | 8.9/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4872 字 阅读 →
论文解读

ROMPAR: Morphological Completion and Demographic Unlearning for Romanian-Accented Speech Recognition

语音识别 | 6.2/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4459 字 阅读 →
论文解读

Scaling Human and G2P Supervision for Robust Phonetic Transcription

语音识别 | 7.6/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4873 字 阅读 →
论文解读

BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM

语音合成 | 9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5563 字 阅读 →
论文解读

Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR

语音识别 | 8.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5961 字 阅读 →
论文解读

Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models

语音识别 | 9.6/10

 · 更新于 2026-09-25 · 约 4 分钟 · 1866 字 阅读 →
论文解读

MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7026 字 阅读 →
论文解读

The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models

语音识别 | 8.2/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5775 字 阅读 →
论文解读

Unsupervised Approaches for Global Prosodic Embedding Extraction

语音合成 | 7.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6086 字 阅读 →
论文解读

Balancing ASR and diarization in end-to-end LLMs for multi-talker speech recognition

语音识别 | 7.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5831 字 阅读 →
论文解读

Positional Encoding in the Context of Memristor-Based Analog Computation for Automatic Speech Recognition

语音识别 | 8/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4178 字 阅读 →
论文解读

PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue

语音合成 | 8.1/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10641 字 阅读 →
论文解读

Evaluating Bias in Phoneme-Based Automatic Speech Recognition Systems: An Analysis of IPA Transcription Models

语音识别 | 8.8/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5179 字 阅读 →
论文解读

Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains

语音识别 | 9.1/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4665 字 阅读 →