论文解读

Improving Active Learning for Melody Estimation by Disentangling Uncertainties

音乐信息检索 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4455 字 阅读 →
论文解读

LLM-Based Post-ASR Error Correction for Disordered Speech

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5031 字 阅读 →
论文解读

Low-Resource Speech-Based Early Alzheimers Detection via Cross-Lingual and Few-Shot Transfer Learning

语音生物标志物 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4145 字 阅读 →
论文解读

Monitoring exposure-length variations in submarine power cables using distributed fiber-optic sensing

音频事件检测 | 6.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3073 字 阅读 →
论文解读

Multimodal Fusion-Based IPCLIP Network for Mixed Reality Surgical Assistance

多模态模型 | 6.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3263 字 阅读 →
论文解读

QFOCUS: Controllable Synthesis for Automated Speech Stress Editing to Deliver Human-Like Emphatic Intent

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 4 分钟 · 1766 字 阅读 →
论文解读

Separate this, and all of these Things Around It: Music Source Separation Via Hyperellipsoidal Queries

音乐分离 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4838 字 阅读 →
论文解读

Stress Prediction from Temporal Emotion Trajectories in Clinical Patient-Physician Conversations

语音情感识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4349 字 阅读 →
论文解读

Synthetic Data Domain Adaptation for ASR via LLM-Based Text and Phonetic Respelling Augmentation

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4965 字 阅读 →
论文解读

Variational Low-Rank Adaptation for Personalized Impaired Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5297 字 阅读 →
论文解读

Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments

音乐生成 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4910 字 阅读 →
论文解读

ICLAD: In-Context Learning with Comparison-Guidance for Audio Deepfake Detection

本文针对音频深度伪造检测模型在真实场景(in-the-wild)中泛化能力差的核心问题,提出了一种名为ICLAD的全新范式。该框架利用音频语言模型(ALM)的上下文学习能力,实现了无需训练的快速适应。其核心是创新的**成对比较推理**策略:在离线阶段,引导ALM为每个样本同时生成“真实”和“伪造”的

 · 更新于 2026-09-25 · 约 13 分钟 · 6206 字 阅读 →
论文解读

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation

这篇论文旨在解决语音到语音翻译(S2ST)系统普遍缺失非语言声音(如笑声、哭泣)和情感韵律的问题,这严重限制了跨语言交流的自然度和语用准确性。作者提出了三大贡献:1) 一个**可扩展的表达性数据合成管道**,能自动生成高质量、带情感标注的S2ST训练对,克服了数据稀缺瓶颈;2) **MoVE(混合声

 · 更新于 2026-09-25 · 约 10 分钟 · 4623 字 阅读 →
论文解读

Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models

本文旨在解决非侵入式语音质量评估在标注数据有限场景下的性能瓶颈。作者提出了GatherMOS框架,其核心是将大语言模型(如GPT-5)作为一个元评估器,通过精心设计的文本提示,融合多类异构信号:包括手

 · 更新于 2026-09-25 · 约 10 分钟 · 4519 字 阅读 →
论文解读

SpeakerRPL v2: Robust Open-set Speaker Identification through Enhanced Few-shot Foundation Tuning and Model Fusion

本文旨在解决开放集说话人识别中的鲁棒性问题,即系统在仅有少量目标说话人注册样本的情况下,需同时准确识别已知说话人并可靠拒识未知说话人。作者在先前SpeakerRPL V1框架基础上提出了三项关键改进:

 · 更新于 2026-09-25 · 约 10 分钟 · 4758 字 阅读 →