论文解读

CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

音视频生成 | 7.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5209 字 阅读 →
论文解读

SAGE: Switch-Aware EEG-Guided Soft Gating for Target Speaker Extraction with In-Trial Switching

语音分离 | 7.3/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4976 字 阅读 →
论文解读

Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds

音频生成 | 7.0/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8583 字 阅读 →
论文解读

SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

语音识别 | 8.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5919 字 阅读 →
论文解读

Sounding Canvas: Embedding Algorithms in Networked, Sensorial Sound Art

RNN | 5.7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5389 字 阅读 →
论文解读

Uncertainty-Aware Crossmodal Fusion for Classification of Animal Behavior

音频分类 | 6.3/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7297 字 阅读 →
论文解读

DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs

音视频语音识别 | 7.4/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9385 字 阅读 →
论文解读

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

音视频理解 | 7.9/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7441 字 阅读 →
论文解读

Technological Advances in Detecting and Managing Cognitive Impairment in Older Adults: Trends, Challenges, and Future Directions

医疗音频 | 4.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6862 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-08-03

共分析 16 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 46 分钟 · 22753 字 阅读 →
论文解读

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

音视频理解 | 8.3/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5406 字 阅读 →
论文解读

Digital Harf: A Clinically Integrated Multimodal AI System for Pervasive Arabic Speech and Language Therapy

语音交互 | 5.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5532 字 阅读 →
论文解读

WeSep: A Modular and Cue-Composable Framework for Target Speaker Extraction

语音分离 | 7.4/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7245 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-31

共分析 16 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 53 分钟 · 26338 字 阅读 →
论文解读

Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

音频伪造检测 | 7.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7402 字 阅读 →
论文解读

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

音乐理解 | 6.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7251 字 阅读 →
论文解读

ThinkOmni: A Reasoning-Driven Omni-Modal LLM Framework for Audio Forgery Detection and Localization

音频伪造检测 | 8.3/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7713 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-30

共分析 19 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 60 分钟 · 30047 字 阅读 →
论文解读

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

多模态模型 | 8.3/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5386 字 阅读 →
论文解读

CARE: A Multimodal Corpus for Studying Speech and Non-Verbal Communication Across Multiple Medical Conditions

音视频理解 | 7.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8510 字 阅读 →