每日研究速递

语音/音乐/音频论文速递 2026-05-25

共分析 19 篇语音/AI 论文

 · 更新于 2026-10-01 · 约 52 分钟 · 25789 字 阅读 →
论文解读

A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation

A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation

 · 更新于 2026-10-01 · 约 1 分钟 · 36 字 阅读 →
论文解读

Abstraction Induces the Brain Alignment of Language and Speech Models

** | 8.0/10

 · 更新于 2026-10-01 · 约 1 分钟 · 90 字 阅读 →
论文解读

Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models

Acoustic Interference: A New Paradigm Weaponizing Acoustic Latent Semantic for Universal Jailbreak against Large Audio Language Models

 · 更新于 2026-10-01 · 约 1 分钟 · 43 字 阅读 →
论文解读

ADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools — From Consensus Learning to Ambiguity-Driven Emotion Reasoning

ADEPT: RL-Aligned Agentic Decoding of Emotion via Evidence Probing Tools — From Consensus Learning to Ambiguity-Driven Emotion Reasoning

 · 更新于 2026-10-01 · 约 1 分钟 · 44 字 阅读 →
论文解读

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

 · 更新于 2026-10-01 · 约 1 分钟 · 37 字 阅读 →
论文解读

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

 · 更新于 2026-10-01 · 约 1 分钟 · 34 字 阅读 →
论文解读

Alethia: a Foundational Encoder for Voice Deepfakes

Alethia: a Foundational Encoder for Voice Deepfakes

 · 更新于 2026-10-01 · 约 1 分钟 · 33 字 阅读 →
论文解读

Any-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

Any-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

 · 更新于 2026-10-01 · 约 1 分钟 · 36 字 阅读 →
论文解读

Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses

Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses

 · 更新于 2026-10-01 · 约 1 分钟 · 40 字 阅读 →
论文解读

AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing

AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing

 · 更新于 2026-10-01 · 约 1 分钟 · 36 字 阅读 →
论文解读

AudioMosaic: Contrastive Masked Audio Representation Learning

AudioMosaic: Contrastive Masked Audio Representation Learning

 · 更新于 2026-10-01 · 约 1 分钟 · 32 字 阅读 →
论文解读

AuTAgent: A Reinforcement Learning Framework for Tool-Augmented Audio Reasoning

AuTAgent: A Reinforcement Learning Framework for Tool-Augmented Audio Reasoning

 · 更新于 2026-10-01 · 约 1 分钟 · 35 字 阅读 →
论文解读

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

 · 更新于 2026-10-01 · 约 1 分钟 · 36 字 阅读 →
论文解读

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

 · 更新于 2026-10-01 · 约 1 分钟 · 33 字 阅读 →
论文解读

AVTrack: Audio-Visual Speaker Tracking in Complex Scenes

AVTrack: Audio-Visual Speaker Tracking in Complex Scenes

 · 更新于 2026-10-01 · 约 1 分钟 · 33 字 阅读 →
论文解读

BAT: Better Audio Transformer Guided by Convex Gated Probing

BAT: Better Audio Transformer Guided by Convex Gated Probing

 · 更新于 2026-10-01 · 约 1 分钟 · 35 字 阅读 →
论文解读

BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps

BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps

 · 更新于 2026-10-01 · 约 1 分钟 · 36 字 阅读 →
论文解读

Bioacoustic Geolocation: Species Sounds as Geographic Signals

Bioacoustic Geolocation: Species Sounds as Geographic Signals

 · 更新于 2026-10-01 · 约 1 分钟 · 33 字 阅读 →
论文解读

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

 · 更新于 2026-10-01 · 约 1 分钟 · 41 字 阅读 →