论文解读

Selective Augmentation: Improving Universal Automatic Phonetic Transcription via G2P Bootstrapping

语音识别 | 6.0/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3160 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-01

共分析 21 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 70 分钟 · 35064 字 阅读 →
论文解读

SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4115 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-04-30

共分析 25 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 80 分钟 · 39601 字 阅读 →
论文解读

A Bimodal Approach for Detecting Fatigue Using Speech and Personal Assessments in College Students

A Bimodal Approach for Detecting Fatigue Using Speech and Personal Assessments in College Students

 · 更新于 2026-09-06 · 约 8 分钟 · 3846 字 阅读 →
论文解读

AFT: An Exemplar-Free Class Incremental Learning Method for Environmental Sound Classification

音频分类 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4408 字 阅读 →
论文解读

Ara-BEST-RQ: Multi Dialectal Arabic SSL

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4587 字 阅读 →
论文解读

Asynchrony-Aware Decoupled Multimodal Control for Cued Speech Video Generation

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4846 字 阅读 →
论文解读

ATOM: Adaptive Token-Level Optimal Transport Mixup for Speech Translation

语音翻译 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4810 字 阅读 →
论文解读

Bayesian Low-Rank Factorization for Robust Model Adaptation

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4176 字 阅读 →
论文解读

Behind the Scenes: Mechanistic Interpretability of Lora-Adapted Whisper for Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3951 字 阅读 →
论文解读

BEST-RQ-based Self-Supervised Learning for Whisper Domain Adaptation

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5130 字 阅读 →
论文解读

BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4898 字 阅读 →
论文解读

CTC-DID: CTC-Based Arabic Dialect Identification for Streaming Applications

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4764 字 阅读 →
论文解读

DDSC: Dynamic Dual-Signal Curriculum for Data-Efficient Acoustic Scene Classification Under Domain Shift

音频场景分类 | 7.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3676 字 阅读 →
论文解读

Domain-Aware Scheduling for ASR Fine-Tuning

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4493 字 阅读 →
论文解读

Efficient Depression Detection from Speech via Language-Independent Prompt-Driven Reprogramming

语音生物标志物 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4393 字 阅读 →
论文解读

Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4068 字 阅读 →
论文解读

Exploring Fine-Tuning Of Large Audio Language Models For Spoken Language Understanding Under Limited Speech Data

语音理解 | 8.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3549 字 阅读 →
论文解读

Fast-ULCNet: A Fast and Ultra Low Complexity Network for Single-Channel Speech Enhancement

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3534 字 阅读 →