论文解读

Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence

声源定位 | 6.8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6953 字 阅读 →
论文解读

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

音频理解 | 7.9/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7542 字 阅读 →
论文解读

Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions

语音质量评估 | 6.0/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9539 字 阅读 →
论文解读

C^2MOE: Consistency and Complementarity-guided Mixture of Experts for Incomplete Multimodal Emotion Learning

音视频理解 | 4.1/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8343 字 阅读 →
论文解读

Equivariant Music Transformer

音乐生成 | 8.6/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9835 字 阅读 →
论文解读

InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion

音频质量评估 | 6.4/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5422 字 阅读 →
论文解读

Modality Agreement- and Conflict-Aware Prototype Hypergraph Learning for Multimodal Intent Understanding

音频理解 | 5.5/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2267 字 阅读 →
论文解读

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

音视频理解 | 7.7/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8473 字 阅读 →
论文解读

OmniVR: Joint Video-Audio Conditional Generation for Restoring Degraded Historical Films

音视频生成 | 7.7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5645 字 阅读 →
论文解读

Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models

音频理解 | 6.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6090 字 阅读 →
论文解读

Towards Robust Version Identification in the Wild: A Dataset, Benchmark, and Fine-Tuning Study

音乐检索 | 8.0/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7467 字 阅读 →
论文解读

Adaptive Modality Reliability Diagnosis and Restoration for Robust Multimodal Intent Recognition

音视频理解 | 6.4/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4713 字 阅读 →
论文解读

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities

音频生成 | 6.9/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6230 字 阅读 →
论文解读

Band-Count Dense Modal Estimation with Fixed-Frequency Differentiable Resonator Refinement

音频理解 | 5.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5290 字 阅读 →
论文解读

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces

语音合成 | 5.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6366 字 阅读 →
论文解读

Calliphony: A Calligraphy-Driven Interface for Real-Time Generative Music Performance

音乐生成 | 5.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6487 字 阅读 →
论文解读

CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis

自回归模型 | 7.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6968 字 阅读 →
论文解读

Cross-cultural evaluation of taste-sound correspondences in AI-generated music

音乐理解 | 7.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9448 字 阅读 →
论文解读

dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

语音编辑 | 7.4/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7954 字 阅读 →
论文解读

Efficient Audio Enhancement with a Differentiable Psychoacoustic Loss

音频超分辨 | 8.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6540 字 阅读 →