论文解读

Qwen-Audio-3.0-Gen-Preview Technical Report

音频生成 | 5.6/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7123 字 阅读 →
论文解读

Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion

语音转换 | 7.9/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5983 字 阅读 →
论文解读

FillGauss: Fine-Grained Filling-Aware Impact Sound Generation for 3D Gaussian Splatting

音频生成 | 6.2/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6266 字 阅读 →
论文解读

Efficient Text-to-Audio Generation via Pruning

音频生成 | 7.6/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5635 字 阅读 →
论文解读

Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models

音频生成 | 7.2/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7410 字 阅读 →
论文解读

A Production-Oriented Framework for Evaluation of SFX Generation

音频生成 | 6.9/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7493 字 阅读 →
论文解读

BackgroundMellow: A Multi-Modal Cohesive Framework for Narrative-Driven Rich Cinematic Soundscape Generation

音频生成 | 7.4/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7705 字 阅读 →
论文解读

FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation

音频生成 | 8.6/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8443 字 阅读 →
论文解读

WaveNet-Style Guitar Amplifier Model Pruning for Real-Time iOS Deployment

音频生成 | 8.0/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5705 字 阅读 →
论文解读

Structural Bottlenecks on Frequency Representation in End-to-End Audio Models

音频生成 | 7.6/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6425 字 阅读 →
论文解读

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space

音视频生成 | 7.4/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7295 字 阅读 →
论文解读

Unified Audio Intelligence Without Regressing on Text Intelligence

音频交互 | 6.8/10

 · 更新于 2026-09-24 · 约 7 分钟 · 3294 字 阅读 →
论文解读

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

语音合成 | 7.6/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7803 字 阅读 →
论文解读

AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing

音频生成 | 5.8/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7545 字 阅读 →
论文解读

Hearing Without Noticing? Attention-Aware Stealthy Black-Box Adversarial Audio Attacks

语音识别 | 7.6/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6579 字 阅读 →
论文解读

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

音频生成 | 7.9/10

 · 更新于 2026-09-24 · 约 20 分钟 · 9942 字 阅读 →
论文解读

Two-dimensional quantization for geometry-aware audio coding

语音编码 | 7.6/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5851 字 阅读 →
论文解读

Predicting Timbre Traits for Interpretable Assessment of Musical Sound Synthesizers

音频生成 | 6.1/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5084 字 阅读 →
论文解读

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation

语音合成 | 7.9/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7085 字 阅读 →
论文解读

LambdaMark: Semantic Audio Watermarking for Robustness and Radioactivity

音频水印 | 8/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7077 字 阅读 →