论文解读

Robust Spoofed Speech Detection via Temporal Pyramid Modeling

音频深度伪造检测 | 6.7/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6597 字 阅读 →
论文解读

ROMPAR: Morphological Completion and Demographic Unlearning for Romanian-Accented Speech Recognition

语音识别 | 6.2/10

 · 更新于 2026-09-07 · 约 9 分钟 · 4459 字 阅读 →
论文解读

Scaling Human and G2P Supervision for Robust Phonetic Transcription

语音识别 | 7.6/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4873 字 阅读 →
论文解读

SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity

大语言模型 | 7.3/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4882 字 阅读 →
论文解读

Semi-Supervised Speech Confidence Detection using Pseudo-Labelling and Whisper Embeddings

Semi-Supervised Speech Confidence Detection using Pseudo-Labelling and Whisper Embeddings

 · 更新于 2026-09-07 · 约 11 分钟 · 5442 字 阅读 →
论文解读

Spectro-Temporal Interference Confounds Phase Encoding in Spatial Audio Foundation Models

自监督学习 | 7.4/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5251 字 阅读 →
论文解读

SPRI: SVD-Partitioned Residual Initialization for Data-Constrained MoE Upcycling

语音翻译 | 7.6/10

 · 更新于 2026-09-07 · 约 14 分钟 · 6597 字 阅读 →
论文解读

Stabilizing Short Duration Speaker Verification through Neural Re-scoring with Hybrid Enrollment

说话人验证 | 7.9/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5781 字 阅读 →
论文解读

Teacher-Student Structure for Domain Adaptation in Ensemble Audio-Visual Video Deepfake Detection

多模态模型 | 7.4/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5649 字 阅读 →
论文解读

TMASC: Transmasculine Attitude and Speech Corpus

TMASC: Transmasculine Attitude and Speech Corpus

 · 更新于 2026-09-07 · 约 11 分钟 · 5137 字 阅读 →
论文解读

Towards Robust Generative Speech Enhancement Using Vector Quantisation-Based Neural Audio Codec

语音增强 | 5.9/10

 · 更新于 2026-09-07 · 约 18 分钟 · 8860 字 阅读 →
论文解读

TuneJury: An Open Metric for Improving Music Generation Preference Alignment

多模态模型 | 9.7/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6483 字 阅读 →
论文解读

Unified Audio Generation and Editing via Joint Condition Modeling and Progressive Training

音频生成 | 8.7/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5275 字 阅读 →
论文解读

Unifying Acoustic Features and Text with Multimodal LLMs for Neurodegenerative Screening

多模态模型 | 6.2/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6481 字 阅读 →
论文解读

Universal adaptive beamforming: A Bayesian approach

自适应滤波 | 8/10

 · 更新于 2026-09-07 · 约 9 分钟 · 4015 字 阅读 →
论文解读

VoxWatermark: A Large-Scale Benchmark for Audio Watermark Detection under Perturbations

鲁棒性 | 9.4/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4873 字 阅读 →
论文解读

When the Same Musical Knowledge Forgets Differently: A Clean Probe of Pathway-Dependent Forgetting

When the Same Musical Knowledge Forgets Differently: A Clean Probe of Pathway-Dependent Forgetting

 · 更新于 2026-09-07 · 约 11 分钟 · 5348 字 阅读 →
论文解读

XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models

多模态模型 | 8.9/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5565 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-16

共分析 62 篇语音/AI 论文

 · 更新于 2026-09-07 · 约 168 分钟 · 83790 字 阅读 →
论文解读

A Deep Zero-Inflated Model of North Atlantic Right Whale Presence To Support Blue Economy Management in the U.S. East Coast

概率图模型 | 7.6/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5651 字 阅读 →