论文解读

AlignSep: Temporally-Aligned Video-Queried Sound Separation with Flow Matching

语音分离 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5425 字 阅读 →
论文解读

Automatic Stage Lighting Control: Is it a Rule-Driven Process or Generative Task?

音乐生成 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5357 字 阅读 →
论文解读

Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4234 字 阅读 →
论文解读

Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction

音乐生成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4687 字 阅读 →
论文解读

LayerSync: Self-aligning Intermediate Layers

生成模型 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4835 字 阅读 →
论文解读

Toward Complex-Valued Neural Networks for Waveform Generation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4635 字 阅读 →
论文解读

ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and Space

条件生成 | 8.0/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3079 字 阅读 →
论文解读

A Generative-First Neural Audio Autoencoder

音乐生成 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5163 字 阅读 →
论文解读

Adaptive Deterministic Flow Matching for Target Speaker Extraction

目标说话人提取 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5050 字 阅读 →
论文解读

Bleed No More: Generative Interference Reduction for Musical Recordings

音乐源分离 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4891 字 阅读 →
论文解读

Combining Multi-Order Attention and Multi-Resolution Discriminator for High-Fidelity Neural Vocoder

语音合成 | 6.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6236 字 阅读 →
论文解读

Confidence-Based Filtering for Speech Dataset Curation with Generative Speech Enhancement Using Discrete Tokens

语音增强 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4790 字 阅读 →
论文解读

Cutscene Agent: An LLM Agent Framework for Automated 3D Cutscene Generation

生成模型 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5619 字 阅读 →
论文解读

ECSA: Dual-Branch Emotion Compensation for Emotion-Consistent Speaker Anonymization

语音匿名化 | 8.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4486 字 阅读 →
论文解读

EmoTri-RL: Emotion- and Cause-Aware Reinforcement Learning for Multi-Modal Empathetic Dialogue

语音情感识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4826 字 阅读 →
论文解读

Enhanced Generative Machine Listener

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4321 字 阅读 →
论文解读

Etude: Piano Cover Generation with a Three-Stage Approach — Extract, Structuralize, and Decode

音乐生成 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4983 字 阅读 →
论文解读

Gen-SER: When the Generative Model Meets Speech Emotion Recognition

语音情感识别 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3718 字 阅读 →
论文解读

Hanui: Harnessing Distributional Discrepancies for Singing Voice Deepfake Detection

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4470 字 阅读 →
论文解读

HCGAN: Harmonic-Coupled Generative Adversarial Network for Speech Super-Resolution in Low-Bandwidth Scenarios

语音增强 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4284 字 阅读 →