论文解读

Music Restoration via Latent Operator Optimization and Diffusion Model Priors

音频修复 | 7.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6911 字 阅读 →
论文解读

On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs

音频理解 | 6.7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5889 字 阅读 →
论文解读

Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning

音频理解 | 7.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6060 字 阅读 →
论文解读

Rethinking Modality Reliability in Multimodal Sentiment Analysis with Incomplete Observations

音视频理解 | 6.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7260 字 阅读 →
论文解读

Simulation-Based Plate-Reverb Parameter Estimation from a Single Impulse Response

音频理解 | 5.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4638 字 阅读 →
论文解读

Speaker Verification Under Real Classroom Conditions for English Speech

说话人验证 | 5.7/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7612 字 阅读 →
论文解读

Towards More Expressive Spoken LLMs: Fine-Grained Intent Benchmarking and Acoustic-Lexical Decoupled Policy Optimization

语音交互 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6106 字 阅读 →
论文解读

Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling

语音合成 | 6.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6963 字 阅读 →
论文解读

What Is Sonification? Toward a Philosophy of Sonification

音频理解 | 3.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5676 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-08-05

共分析 32 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 93 分钟 · 46258 字 阅读 →
论文解读

Allocation Before Ranking: Decoupled Token Compression for OmniLLMs

音视频理解 | 8.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6640 字 阅读 →
论文解读

Analyzing Speech Condition Effects in Dysarthric ASR: A Layer-wise Probing Study

语音识别 | 5.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6341 字 阅读 →
论文解读

Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation

语音合成 | 7.8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6646 字 阅读 →
论文解读

Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding

语音属性识别 | 6.9/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5046 字 阅读 →
论文解读

DRONEAUDIONET: Noise Suppression for Drone Audition-based Search and Rescue

音频分离 | 7.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5436 字 阅读 →
论文解读

Embodied Empathy: A Multimodal AR and LLM-Powered System for Self-Attachment Psychotherapy with Self-Initiated Humour

音频交互 | 5.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6236 字 阅读 →
论文解读

Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech

语音合成 | 7.2/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5978 字 阅读 →
论文解读

FATE: Frame-Level Audio-Visual Temporal Embedding

音视频理解 | 8.3/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6937 字 阅读 →
论文解读

Gecko: Fast Private Inference via Secure Public Encoder Offloading

迁移学习 | 6.9/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7533 字 阅读 →
论文解读

Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

音频理解 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6085 字 阅读 →