论文解读

SAGA-SR: Semantically and Acoustically Guided Audio Super-Resolution

音频增强 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4383 字 阅读 →
论文解读

Scalable Evaluation for Audio Identification Via Synthetic Latent Fingerprint Generation

音频检索 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3896 字 阅读 →
论文解读

SFM-TTS: Lightweight and Rapid Speech Synthesis with Flexible Shortcut Flow Matching

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5280 字 阅读 →
论文解读

Shortcut Flow Matching for Speech Enhancement: Step-Invariant Flows via Single Stage Training

语音增强 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4631 字 阅读 →
论文解读

Single-Step Controllable Music Bandwidth extension with Flow Matching

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5522 字 阅读 →
论文解读

Stemphonic: All-At-Once Flexible Multi-Stem Music Generation

音乐生成 | 7.7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4942 字 阅读 →
论文解读

StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks

歌唱语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5251 字 阅读 →
论文解读

Symphony Rendering: Midi and Composer-Conditioned Auto Orchestration with Flow-Matching Transformers

音乐生成 | 7.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5626 字 阅读 →
论文解读

Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4678 字 阅读 →
论文解读

TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset Generation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4601 字 阅读 →
论文解读

Towards Real-Time Generative Speech Restoration with Flow-Matching

语音增强 | 6.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4390 字 阅读 →
论文解读

Training Flow Matching Models with Reliable Labels via Self-Purification

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4189 字 阅读 →
论文解读

Universr: Unified and Versatile Audio Super-Resolution Via Vocoder-Free Flow Matching

音频超分辨率 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4168 字 阅读 →
论文解读

V2A-DPO: Omni-Preference Optimization for Video-To-Audio Generation

视频到音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4613 字 阅读 →
论文解读

VoxMorph: Scalable Zero-Shot Voice Identity Morphing via Disentangled Embeddings

语音克隆 | 9.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5242 字 阅读 →
论文解读

MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4497 字 阅读 →
论文解读

Speech Enhancement Based on Drifting Models

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5266 字 阅读 →
论文解读

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6055 字 阅读 →
论文解读

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

音频生成 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5972 字 阅读 →
论文解读

MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3859 字 阅读 →