论文解读

WaveSSM: Multiscale State-Space Models for Non-stationary Signal Attention

音频分类 | 4.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5772 字 阅读 →
论文解读

An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation

语音合成 | 6.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6302 字 阅读 →
论文解读

Beyond Words: Towards Effective Modeling of Non-Verbal Vocalizations in ASR

语音识别 | 6.4/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8399 字 阅读 →
论文解读

From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages

语音识别 | 4.8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6918 字 阅读 →
论文解读

H-SAGE: Holistic Speaker-Aware Guided Experts for MoE-based Multi-Talker ASR

语音识别 | 6.9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5995 字 阅读 →
论文解读

NAVER LABS Europe Submission to the Instruction-following 2026 Short Track

语音翻译 | 6.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7251 字 阅读 →
论文解读

Spatial Speech Perception Systems: A Survey of Sound Source Localization, Directional Enhancement, and Speech Recognition

声源定位 | 4.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5837 字 阅读 →
论文解读

NPUsper: Eliminating Redundant Computation for Real-Time Whisper on Mobile NPUs

语音识别 | 9/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7898 字 阅读 →
论文解读

Adapting Foundation ASR Models to Dysarthric Speech: A Case Study

语音识别 | 6.2/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5802 字 阅读 →
论文解读

BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations

语音识别 | 6.9/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4866 字 阅读 →
论文解读

Building a Multimodal Dataset of Academic Paper for Keyword Extraction

语音识别 | 5.2/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5542 字 阅读 →
论文解读

Building an ASR Solution for Training and Assessing Children's Reading

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5557 字 阅读 →
论文解读

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition

语音识别 | 8.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5117 字 阅读 →
论文解读

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

语音合成 | 7.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6790 字 阅读 →
论文解读

Improving multichannel speech enhancement through accurate room-acoustic simulations

语音增强 | 6.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5624 字 阅读 →
论文解读

LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6067 字 阅读 →
论文解读

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation

语音识别 | 7/10

 · 更新于 2026-09-25 · 约 6 分钟 · 2889 字 阅读 →
论文解读

Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition

语音识别 | 7.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5903 字 阅读 →
论文解读

What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR

语音识别 | 7.3/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12422 字 阅读 →
论文解读

Agent-Computer Observation Interfaces Enable Dynamic Computer Use

语音识别 | 8.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6790 字 阅读 →