论文解读

AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5530 字 阅读 →
论文解读

AlignSep: Temporally-Aligned Video-Queried Sound Separation with Flow Matching

语音分离 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5425 字 阅读 →
论文解读

Aurelius: Relation Aware Text-to-Audio Generation At Scale

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4737 字 阅读 →
论文解读

Continuous Audio Language Models

音频生成 音乐生成 | 9.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5470 字 阅读 →
论文解读

Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4234 字 阅读 →
论文解读

FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6079 字 阅读 →
论文解读

Gogo: Group-wise granularity-ordered codec for stable and efficient speech generation

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4704 字 阅读 →
论文解读

InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

视频生成 | 9.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4843 字 阅读 →
论文解读

JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation

音视频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6841 字 阅读 →
论文解读

LayerSync: Self-aligning Intermediate Layers

生成模型 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4835 字 阅读 →
论文解读

NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching

多模态模型 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5125 字 阅读 →
论文解读

Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis

音频分类 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 5010 字 阅读 →
论文解读

Scaling Speech Tokenizers with Diffusion Autoencoders

语音分词 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4954 字 阅读 →
论文解读

Stable Video Infinity: Infinite-Length Video Generation with Error Recycling

视频生成 | 8.8/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4252 字 阅读 →
论文解读

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

音频生成 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4891 字 阅读 →
论文解读

Unified Multi-Modal Interactive and Reactive 3D Motion Generation via Rectified Flow

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4970 字 阅读 →
论文解读

JaiTTS: A Thai Voice Cloning Model

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2500 字 阅读 →
论文解读

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

音频生成 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5965 字 阅读 →
论文解读

Adaptive Deterministic Flow Matching for Target Speaker Extraction

目标说话人提取 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5050 字 阅读 →
论文解读

AnyAccomp: Generalizable Accompaniment Generation Via Quantized Melodic Bottleneck

音乐生成 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4201 字 阅读 →