论文解读

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

音视频生成 | 6.6/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6714 字 阅读 →
论文解读

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation

知识蒸馏 | 10/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5503 字 阅读 →
论文解读

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6593 字 阅读 →
论文解读

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

语音合成 | 7.3/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2255 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-01

共分析 35 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 102 分钟 · 50963 字 阅读 →
论文解读

CTC-Seeded Token Edit Refinement for Non-Autoregressive Speech Recognition

语音识别 | 7.7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5815 字 阅读 →
论文解读

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

音乐生成 | 9.4/10

 · 更新于 2026-09-25 · 约 4 分钟 · 1658 字 阅读 →
论文解读

SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L Dataset

语音合成 | 8.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6507 字 阅读 →
论文解读

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

语音合成 | 6/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4968 字 阅读 →
论文解读

UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating

扩散模型 | 8.9/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7367 字 阅读 →
论文解读

Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS

语音合成 | 7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5086 字 阅读 →
论文解读

A Variational-Flow Analysis of StoRM under Noise-Power Mismatch

语音增强 | 4.4/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4515 字 阅读 →
论文解读

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

语音增强 | 8.1/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4879 字 阅读 →
论文解读

Bridging Self-Supervised Learning and Speech Enhancement: A Wav2Vec2-Conditioned Framework

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5269 字 阅读 →
论文解读

ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion

语音转换 | 6.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5546 字 阅读 →
论文解读

Co-policy: Responsive Human-Robot Co-Creation for Musical Performances

音乐生成 | 8.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6648 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-22

共分析 1 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 4 分钟 · 1684 字 阅读 →
论文解读

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

语音合成 | 7.7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5246 字 阅读 →
论文解读

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

多模态模型 | 8.1/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5119 字 阅读 →
论文解读

Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation

语音合成 | 7.9/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6654 字 阅读 →