论文解读

UAT: Unified Audio-Text Diffusion for Audio Generation, Editing, and Captioning

音频生成 | 8.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6372 字 阅读 →
论文解读

Cosmos 3: Omnimodal World Models for Physical AI

音频生成 | 10/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8607 字 阅读 →
论文解读

Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation

音频生成 | 7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6522 字 阅读 →
论文解读

SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling

音乐生成 | 8.6/10

 · 更新于 2026-09-25 · 约 26 分钟 · 12623 字 阅读 →
论文解读

UniVocal: Unified Speech-Singing Code-Switching Synthesis

语音合成 | 8.9/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2097 字 阅读 →
论文解读

Sound effects in media:A comparative analysis of recorded and synthetic samples in live-action and animation

音频生成 | 5.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6218 字 阅读 →
论文解读

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion

语音合成 | 8.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7971 字 阅读 →
论文解读

Benchmarking Single-Factor Physical Video-to-Audio Generation

音频生成 | 9/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6510 字 阅读 →
论文解读

Native Audio-Visual Alignment for Generation

音频生成 | 7.8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6977 字 阅读 →
论文解读

Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text

音频生成 | 8.6/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7601 字 阅读 →
论文解读

LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation

语音合成 | 7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6912 字 阅读 →
论文解读

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

音频生成 | 7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4951 字 阅读 →
论文解读

CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation

音频生成 | 8.7/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7201 字 阅读 →
论文解读

CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation

音频生成 | 6/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7631 字 阅读 →
论文解读

Stable Audio 3

音频生成 | 6.8/10

 · 更新于 2026-09-25 · 约 23 分钟 · 11116 字 阅读 →
论文解读

Taming Audio VAEs via Target-KL Regularization

音频生成 语音合成 | 6.7/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7824 字 阅读 →
论文解读

WavFlow: Audio Generation in Waveform Space

音频生成 | 6.7/10

 · 更新于 2026-09-25 · 约 24 分钟 · 11712 字 阅读 →
论文解读

Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis

音频生成 | 6.8/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8349 字 阅读 →
论文解读

FSD50K-Solo: Automated Curation of Single-Source Sound Events

数据清洗 | 5.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6839 字 阅读 →
论文解读

Seconds-Aligned PCA-DAC Latent Diffusion for Symbolic-to-Audio Drum Rendering

音频生成 | 7.0/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10175 字 阅读 →