论文解读

FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching

视频生成 | 4.9/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7098 字 阅读 →
论文解读

PresentAgent-2: Towards Generalist Multimodal Presentation Agents

生成模型 | 6.5/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7708 字 阅读 →
会议任务专题

ICLR 2026 - 视频生成

共 2 篇 ICLR 2026 视频生成 方向论文

 · 更新于 2026-09-24 · 约 6 分钟 · 2744 字 阅读 →
论文解读

InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

视频生成 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4928 字 阅读 →
论文解读

Stable Video Infinity: Infinite-Length Video Generation with Error Recycling

视频生成 | 7.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4282 字 阅读 →
论文解读

InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

视频生成 | 9.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4843 字 阅读 →
论文解读

Stable Video Infinity: Infinite-Length Video Generation with Error Recycling

视频生成 | 8.8/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4252 字 阅读 →
论文解读

Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers

音频生成 | 7.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5245 字 阅读 →
会议任务专题

ICASSP 2026 - 视频生成

共 2 篇 ICASSP 2026 视频生成 方向论文

 · 更新于 2026-09-24 · 约 6 分钟 · 2599 字 阅读 →
论文解读

MirrorTalk: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion Control

语音合成 | 7.0/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3679 字 阅读 →
论文解读

StyHarmo: Efficient Style-Specific Video Generation with Music Synchronization

视频生成 | 6.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5066 字 阅读 →
论文解读

VT-Heads: Voice Cloning and Talking Head Generation from Text Based on V-DiT

视频生成 | 6.5/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3563 字 阅读 →
论文解读

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation

1. **问题**:现有视频扩散模型在生成人机交互(HOI)视频时,常出现手/脸结构崩溃和人机物理穿透等问题,根源在于模型缺乏对3D空间关系和交互结构的理解。 2. **方法核心**:提出CoInteract框架,核心是“空间结构化协同生成”范式。在一个共享的DiT骨干中联合训练RGB外观流和辅助的

 · 更新于 2026-09-24 · 约 11 分钟 · 5017 字 阅读 →