论文解读

AudioNoisePrints: Model-free audio watermarking using spatial correlation in flow matching TTS

音频水印 | 7.9/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4737 字 阅读 →
论文解读

AudioWorldSim: Realistic Binaural Audio Datasets For World Models

空间音频 | 9.7/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4341 字 阅读 →
论文解读

FlowSep 2: Self-Supervised Flow Matching for Language-Queried Audio Source Separation

音频分离 | 9.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4945 字 阅读 →
论文解读

Geometric Iterative Retrieval for Neural Audio Codec Resynthesis

音频编码 | 7.6/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5287 字 阅读 →
论文解读

DynaForcing: Overcoming Dynamic Collapse in Self-Forcing Distillation for Streaming Avatar Generation

音视频生成 | 8.0/10

 · 更新于 2026-09-24 · 约 7 分钟 · 3309 字 阅读 →
论文解读

FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations

语音合成 | 9.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4189 字 阅读 →
论文解读

Numerical and perceptual validity of synthetic Head-Related Transfer Functions at scale

声源定位 | 7.6/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8223 字 阅读 →
论文解读

Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences

音乐生成 | 5.2/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7478 字 阅读 →
论文解读

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping

音乐生成 | 6.8/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8231 字 阅读 →
论文解读

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies

音频生成 | 6.5/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8964 字 阅读 →
论文解读

Beyond Reconstruction: Full-Context Generative DiT for Music Generation

音乐生成 | 7.5/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7890 字 阅读 →
论文解读

Beyond Call and Response: Modelling Reciprocal Coordination in Human-AI Vocal Ensembles

歌唱生成 | 5.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5433 字 阅读 →
论文解读

PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement

语音增强 | 8.1/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7802 字 阅读 →
论文解读

AnyBand: Unified Multi-Bandwidth Speech Extension via Frequency-Aware In-Context Spectral Infilling

语音超分 | 5.9/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5441 字 阅读 →
论文解读

Separate-and-Detect: Unified Drum Transcription and Stem Generation via Latent Diffusion

音乐转录 | 6.8/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5561 字 阅读 →
论文解读

MusiChat: Vibe Composing for Music Creation

音乐生成 | 5.8/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6218 字 阅读 →
论文解读

MemNMF: Memory-Augmented NMF on LPC Spectra for Anomalous Sound Detection

音频事件检测 | 5.3/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6236 字 阅读 →
论文解读

CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses

语音合成 | 5.3/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8594 字 阅读 →
论文解读

ITGPT: A Transformer Based Architecture for the Generation of Dance Dance Revolution and In the Groove Charts

音乐生成 | 6.5/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7824 字 阅读 →
论文解读

Bring Music The Horizon: Music-Driven 360^\circ Video Generation

音视频生成 | 5.3/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4986 字 阅读 →