论文解读

Omni-Customizer: End-to-End MultiModal Customization for Joint Audio-Video Generation

音视频 | 7.3/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9061 字 阅读 →
论文解读

SAME: A Semantically-Aligned Music Autoencoder

音频编码 | 8.5/10

 · 更新于 2026-09-24 · 约 26 分钟 · 13010 字 阅读 →
论文解读

PresentAgent-2: Towards Generalist Multimodal Presentation Agents

生成模型 | 6.5/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7708 字 阅读 →
论文解读

Drum Synthesis from Expressive Drum Grids via Neural Audio Codecs

音乐生成 | 4.0/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9279 字 阅读 →
论文解读

PoDAR: Power-Disentangled Audio Representation for Generative Modeling

语音合成 | 7.3/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8991 字 阅读 →
论文解读

Do Joint Audio-Video Generation Models Understand Physics?

Do Joint Audio-Video Generation Models Understand Physics?

 · 更新于 2026-09-24 · 约 18 分钟 · 8531 字 阅读 →
论文解读

Audio-Visual Intelligence in Large Foundation Models

跨模态 | 8.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4580 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-09

共分析 3 篇语音/AI 论文

 · 更新于 2026-09-24 · 约 10 分钟 · 4857 字 阅读 →
论文解读

RenCon 2025: Revival of the Expressive Performance Rendering Competition

音乐生成 | 7.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4621 字 阅读 →
论文解读

Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement

语音增强 | 7.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5050 字 阅读 →
论文解读

Stage Light is Sequence$^2$: Multi-Light Control via Imitation Learning

音乐信息检索 | 7.5/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7610 字 阅读 →
论文解读

MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech

音频安全 | 7.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5419 字 阅读 →
论文解读

AUHead: Realistic Emotional Talking Head Generation via Action Units Control

生成模型 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4862 字 阅读 →
论文解读

Confident and Adaptive Generative Speech Recognition via Risk Control

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4401 字 阅读 →
论文解读

DiVeQ: Differentiable Vector Quantization Using the Reparameterization Trick

生成模型 | 8.0/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5814 字 阅读 →
会议任务专题

ICLR 2026 - 生成模型

共 2 篇 ICLR 2026 生成模型 方向论文

 · 更新于 2026-09-24 · 约 8 分钟 · 3559 字 阅读 →
论文解读

LayerSync: Self-aligning Intermediate Layers

音频生成 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4057 字 阅读 →
论文解读

Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers

音视频 | 8.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5201 字 阅读 →
论文解读

Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation

声源定位 | 8.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4493 字 阅读 →
论文解读

A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers

生成模型 | 6.5/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3762 字 阅读 →