论文解读

Fair Cognitive Impairment Detection Through Unlearning

多模态模型 | 7.7/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5572 字 阅读 →
论文解读

GRIDEX: Grid-Grounded Forensic Explanations for Deepfake Spectrogram Analysis

语音合成 | 8.6/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7676 字 阅读 →
论文解读

Learning Robust Pair Confidence for Multimodal Emotion-Cause Pair Extraction

多模态模型 | 7.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6808 字 阅读 →
论文解读

Mitigating Scoring Errors and Compensating for Nonverbal Subtests in Speech-Based Dementia Assessment

多模态模型 | 8/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6686 字 阅读 →
论文解读

Native Active Perception as Reasoning for Omni-Modal Understanding

语音识别 | 9.1/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6616 字 阅读 →
论文解读

Risk Stratification for ICU Delirium using Pervasive Ambient Sensing Information

多模态模型 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4563 字 阅读 →
论文解读

ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection

强化学习 | 6.3/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4793 字 阅读 →
论文解读

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs

多模态模型 | 8.8/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5845 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-18

共分析 36 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 105 分钟 · 52107 字 阅读 →
论文解读

A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models

多模态模型 | 6.6/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5184 字 阅读 →
论文解读

Grounding Spoken LLMs in Multi-Speaker Audio via Diarization Conditioning

语音识别 | 8.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5868 字 阅读 →
论文解读

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

语音合成 | 7.7/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4464 字 阅读 →
论文解读

OlfactProfile: Profile-Conditioned Odor Prediction from Audiovisual Content

多模态模型 | 5.6/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5084 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-17

共分析 35 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 94 分钟 · 46631 字 阅读 →
论文解读

Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models

音频分类 | 8.3/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5504 字 阅读 →
论文解读

Bridging the SEA Gap: An Initial Benchmark for Neural Audio Codec-Synthesized Speech Deepfakes in South-East Asian Languages

语音合成 | 8.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5381 字 阅读 →
论文解读

EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning

音频问答 | 6.1/10

 · 更新于 2026-09-06 · 约 22 分钟 · 10753 字 阅读 →
论文解读

MAF: Multimodal Adaptive Few-shot Prompting for Sentiment Analysis with MLLMs

多模态模型 | 5.9/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6343 字 阅读 →
论文解读

MUNI: Multimodal Unified Latent Diffusion for Coherent Any-to-Any Generation

语音生成 | 6.9/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5373 字 阅读 →
论文解读

MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild

语音对话系统 | 7.8/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6606 字 阅读 →