论文解读

AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4618 字 阅读 →
论文解读

Alethia: A Foundational Encoder for Voice Deepfakes

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4189 字 阅读 →
论文解读

AlignSep: Temporally-Aligned Video-Queried Sound Separation with Flow Matching

音频分离 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4519 字 阅读 →
论文解读

Aurelius: Relation Aware Text-to-Audio Generation At Scale

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4141 字 阅读 →
论文解读

Continuous Audio Language Models

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5685 字 阅读 →
论文解读

FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates

语音合成 | 9.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5627 字 阅读 →
论文解读

Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6318 字 阅读 →
论文解读

FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows

跨模态生成 | 9.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5159 字 阅读 →
论文解读

Gogo: Group-wise granularity-ordered codec for stable and efficient speech generation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5137 字 阅读 →
论文解读

Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech Synthesis

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9658 字 阅读 →
论文解读

JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation

音视频 | 9.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4810 字 阅读 →
论文解读

MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5854 字 阅读 →
论文解读

NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching

多模态模型 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5387 字 阅读 →
论文解读

PrismAudio: Decomposed Chain-of-Thought and Multi-dimensional Rewards for Video-to-Audio Generation

音频生成 | 9.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4846 字 阅读 →
论文解读

Resp-Agent: An Agent-Based System for Multimodal Respiratory Sound Generation and Disease Diagnosis

音频分类 | 9.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6625 字 阅读 →
论文解读

Scaling Speech Tokenizers with Diffusion Autoencoders

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3677 字 阅读 →
论文解读

Stable Video Infinity: Infinite-Length Video Generation with Error Recycling

视频生成 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4282 字 阅读 →
论文解读

Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers

音视频 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5201 字 阅读 →
论文解读

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 24 分钟 · 12017 字 阅读 →
论文解读

Unified Multi-Modal Interactive and Reactive 3D Motion Generation via Rectified Flow

动作生成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4802 字 阅读 →