论文解读

Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens

语音识别 | 6.3/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8991 字 阅读 →
论文解读

DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues

语音交互 | 6.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5720 字 阅读 →
论文解读

Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement

音频分类 | 6.8/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7590 字 阅读 →
论文解读

Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

音频伪造检测 | 7.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7402 字 阅读 →
论文解读

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

音频字幕生成 | 6.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6369 字 阅读 →
论文解读

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

音频交互 | 7.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6613 字 阅读 →
论文解读

Qwen-Audio-3.0-Gen-Preview Technical Report

音频生成 | 5.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7123 字 阅读 →
论文解读

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

音乐理解 | 6.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7251 字 阅读 →
论文解读

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

音视频生成 | 6.7/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7788 字 阅读 →
论文解读

Unfolded Recursive Expectation-Maximization Neural Network For Speaker Tracking

声源定位 | 5.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5982 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-30

共分析 19 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 60 分钟 · 30047 字 阅读 →
论文解读

A Cross-lingual Comparison of Human and Classification Model Entrainment Behavior in Code-switched Speech Settings

Transformer | 5.1/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7784 字 阅读 →
论文解读

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

多模态模型 | 8.3/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5386 字 阅读 →
论文解读

Depression Markers in Speech: An Approach based on Tract Variables Dynamics

语音属性识别 | 5.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6197 字 阅读 →
论文解读

Device Invariance using Domain Adaptation on Acoustic Scene Classification

领域适应 | 4.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6363 字 阅读 →
论文解读

DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment

音视频理解 | 4.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8175 字 阅读 →
论文解读

Evaluation of forced alignment of code-mixed speech: the case of Hindi-English

语音识别 | 5.4/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4432 字 阅读 →
论文解读

Finding the noise: Zero-shot AI Music Detection

无监督学习 | 5.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6037 字 阅读 →
论文解读

From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding

音频理解 | 5.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6416 字 阅读 →
论文解读

Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception

语音交互 | 7.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7535 字 阅读 →