每日研究速递

语音/音乐/音频论文速递 2026-06-08

共分析 38 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 108 分钟 · 53770 字 阅读 →
论文解读

Age-Aware Adapter Tuning for Children's Speech Recognition

语音识别 | 8.4/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5531 字 阅读 →
论文解读

An Ultra-Low-Bitrate Neural Speech Codec with Plain-to-Pseudo Synergistic Vector Quantization

语音合成 | 7.7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5324 字 阅读 →
论文解读

Automatic Labelling of Speech Translation Errors

语音识别 | 6.1/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4750 字 阅读 →
论文解读

Beyond Generative Decoding: Discriminative Hidden-State Readout from a Native Omni-Modal LLM for Multimodal Sentiment Analysis

多模态模型 | 5.3/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6789 字 阅读 →
论文解读

CoSTA: Cognitive-State-Conditioned TTS Data Augmentation Using ASR Transcripts for Alzheimer's Disease Detection

语音合成 | 6.5/10

 · 更新于 2026-09-06 · 约 4 分钟 · 1509 字 阅读 →
论文解读

Exploring LLMs for South Asian Music Understanding and Generation

音乐生成 | 7.7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5025 字 阅读 →
论文解读

FiLM-Based Speaker Conditioning of a SpeechLLM for Pathological Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7310 字 阅读 →
论文解读

Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition

语音识别 | 6.9/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4724 字 阅读 →
论文解读

Multilingual Detection of Alzheimer's Disease from Speech: A Cross-Linguistic Transfer Learning Approach

迁移学习 | 5.7/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4489 字 阅读 →
论文解读

Revisiting Lexicon Evaluation in Unsupervised Word Discovery

语音识别 | 1/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5346 字 阅读 →
论文解读

Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs

语音识别 | 5.9/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5559 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-05

共分析 47 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 127 分钟 · 63374 字 阅读 →
论文解读

Masked Wavelet Scattering Transform Neural Field for Sound Field Reconstruction

音频质量评估 | 6.7/10

 · 更新于 2026-09-06 · 约 28 分钟 · 13534 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-04

共分析 22 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 60 分钟 · 29896 字 阅读 →
论文解读

A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026

语音翻译 | 6.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6364 字 阅读 →
论文解读

BaltiVoice: A Speech Corpus and Fine-tuned Whisper ASR System for the Balti Language

语音识别 | 7.8/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4606 字 阅读 →
论文解读

Benchmarking Speech-to-Speech Translation Models

语音合成 | 8.7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5395 字 阅读 →
论文解读

Efficient ASR Training with Conversations that Never Happened

语音识别 | 8/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6552 字 阅读 →
论文解读

FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations

语音识别 | 8.1/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6820 字 阅读 →