论文解读

Segment-Level Mandarin Chinese Speech-Based Cognitive Impairment Detection via an Autoencoder with Contrastive Learning

对比学习 | 6.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5574 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-19

共分析 40 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 110 分钟 · 54842 字 阅读 →
论文解读

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

语音识别 | 9.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5563 字 阅读 →
论文解读

Low-resource Language Discrimination Towards Chinese Dialects with Transfer learning and Data Augmentation

语音识别 | 5.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5615 字 阅读 →
论文解读

Montreal Forced Aligner and the state of speech-to-text alignment in 2026

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6323 字 阅读 →
论文解读

NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization

声源定位 | 7.8/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5653 字 阅读 →
论文解读

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs

语音合成 | 7.4/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7228 字 阅读 →
论文解读

Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6503 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-18

共分析 36 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 105 分钟 · 52107 字 阅读 →
论文解读

An Analysis of the Effectiveness of Synthetic Speech Data for ASR Fine-tuning in Selected Indic Languages

语音识别 | 5.7/10

 · 更新于 2026-09-06 · 约 25 分钟 · 12144 字 阅读 →
论文解读

Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4484 字 阅读 →
论文解读

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

语音合成 | 7.7/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4464 字 阅读 →
论文解读

When Multiple Scripts Matter: Evaluating ASR in Clinical Settings

语音识别 | 9.1/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6513 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-17

共分析 35 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 94 分钟 · 46631 字 阅读 →
论文解读

An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis

语音合成 | 8.2/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4584 字 阅读 →
论文解读

ArtBoost: Synthetic Articulatory Data Augmentation for Acoustic-to-Articulatory Inversion

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5884 字 阅读 →
论文解读

Beyond Classification: A Cough Regression Benchmark for Respiratory Acoustic Foundation Models

音频事件检测 | 6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6382 字 阅读 →
论文解读

Bridging the SEA Gap: An Initial Benchmark for Neural Audio Codec-Synthesized Speech Deepfakes in South-East Asian Languages

语音合成 | 8.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5381 字 阅读 →
论文解读

Confidence Score Guided Incremental and Speaker Adaptive Pseudo-Labeling for Semi-Supervised Elderly Speech Recognition

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6027 字 阅读 →
论文解读

CraBERT: Efficient Phoneme Encoder Pre-Training via Cascade Fusion of Subword Representations for Text-to-Speech

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5796 字 阅读 →