论文解读

Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration

音乐生成 | 7.5/10

 · 更新于 2026-10-01 · 约 18 分钟 · 8830 字 阅读 →
论文解读

Probing Cross-modal Information Hubs in Audio-Visual LLMs

模型分析 | 6.5/10

 · 更新于 2026-10-01 · 约 19 分钟 · 9363 字 阅读 →
论文解读

RADAR Challenge 2026: Robust Audio Deepfake Recognition under Media Transformations

音频深度伪造检测 | 6.0/10

 · 更新于 2026-10-01 · 约 16 分钟 · 7522 字 阅读 →
论文解读

Reducing Linguistic Hallucination in LM-Based Speech Enhancement via Noise-Invariant Acoustic-Semantic Distillation

语音增强 | 7.2/10

 · 更新于 2026-10-01 · 约 21 分钟 · 10260 字 阅读 →
论文解读

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems

音色迁移 | 5.5/10

 · 更新于 2026-10-01 · 约 18 分钟 · 8892 字 阅读 →
论文解读

Responsible Benchmarking of Fairness for Automatic Speech Recognition

语音识别 | 5.0/10

 · 更新于 2026-10-01 · 约 16 分钟 · 7970 字 阅读 →
论文解读

Rethinking Entropy Minimization in Test-Time Adaptation for Autoregressive Models

语音识别 | 6.0/10

 · 更新于 2026-10-01 · 约 19 分钟 · 9094 字 阅读 →
论文解读

Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought

音视频问答 | 6.0/10

 · 更新于 2026-10-01 · 约 20 分钟 · 9627 字 阅读 →
论文解读

SF-Flow: Sound field magnitude estimation via flow matching guided by sparse measurements

空间音频 | 6.8/10

 · 更新于 2026-10-01 · 约 16 分钟 · 7898 字 阅读 →
论文解读

ShipEcho -- An Interactive Tool for Global Mapping of Underwater Radiated Noise from Vessels

水下声学 | 6.0/10

 · 更新于 2026-10-01 · 约 15 分钟 · 7190 字 阅读 →
论文解读

Single-Microphone Audio Point Source Discriminative Localization From Reverberation Late Tail Estimation

说话人分离 | 5.0/10

 · 更新于 2026-10-01 · 约 17 分钟 · 8189 字 阅读 →
论文解读

Speech-based Psychological Crisis Assessment using LLMs

语音情感识别 | 5.8/10

 · 更新于 2026-10-01 · 约 17 分钟 · 8451 字 阅读 →
论文解读

Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Models

世界模型 | 5.0/10

 · 更新于 2026-10-01 · 约 10 分钟 · 4662 字 阅读 →
论文解读

Towards Trustworthy Audio Deepfake Detection: A Systematic Framework for Diagnosing and Mitigating Gender Bias

音频深度伪造检测 | 6.5/10

 · 更新于 2026-10-01 · 约 18 分钟 · 8536 字 阅读 →
论文解读

Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation

音视频生成 | 6.5/10

 · 更新于 2026-10-01 · 约 19 分钟 · 9059 字 阅读 →
论文解读

Voice Biomarkers for Depression and Anxiety

语音生物标志物 | 1.0/10

 · 更新于 2026-10-01 · 约 10 分钟 · 4782 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-12

共分析 39 篇语音/AI 论文

 · 更新于 2026-10-01 · 约 157 分钟 · 78445 字 阅读 →
论文解读

A Decomposed Retrieval-Edit-Rerank Framework for Chord Generation

A Decomposed Retrieval-Edit-Rerank Framework for Chord Generation

 · 更新于 2026-10-01 · 约 15 分钟 · 7086 字 阅读 →
论文解读

Adaptive Regularization for Sparsity Control in Bregman-Based Optimizers

Adaptive Regularization for Sparsity Control in Bregman-Based Optimizers

 · 更新于 2026-10-01 · 约 16 分钟 · 7952 字 阅读 →
论文解读

Anisotropic Modality Align

Anisotropic Modality Align

 · 更新于 2026-10-01 · 约 16 分钟 · 7980 字 阅读 →