论文解读

UALM: Unified Audio Language Model for Understanding, Generation and Reasoning

音频生成 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4691 字 阅读 →
论文解读

Unified Multi-Modal Interactive and Reactive 3D Motion Generation via Rectified Flow

音频生成 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4970 字 阅读 →
论文解读

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Video

多模态模型 | 7.0/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3342 字 阅读 →
论文解读

WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

音频检索 视频检索 | 8.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5230 字 阅读 →
论文解读

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

基准测试 | 8.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4207 字 阅读 →
论文解读

XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models

基准测试 | 9.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4501 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-02

共分析 4 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 15 分钟 · 7026 字 阅读 →
论文解读

BUT System Description for CHiME-9 MCoRec Challenge

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5308 字 阅读 →
论文解读

Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge

语音对话系统 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3578 字 阅读 →
论文解读

Mapping the Methodological Space of Classroom Interaction Research: Scale, Duration, and Modality in an Age of AI

模型评估 | 6.0/10

 · 更新于 2026-09-06 · 约 6 分钟 · 2848 字 阅读 →
论文解读

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

语音对话系统 | 8.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6664 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-01

共分析 21 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 70 分钟 · 35064 字 阅读 →
论文解读

A Bimodal Approach for Detecting Fatigue Using Speech and Personal Assessments in College Students

A Bimodal Approach for Detecting Fatigue Using Speech and Personal Assessments in College Students

 · 更新于 2026-09-06 · 约 8 分钟 · 3846 字 阅读 →
论文解读

A Dynamic Gated Cross-Attention Framework for Audio-Text Apparent Personality Analysis

音频分类 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4050 字 阅读 →
论文解读

ACIR-MACL: Effective Multimodal Sentiment Analysis via Attention-Based Causal Intervention Regularization and Multi-Aspect Contrastive Learning

情感分析 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5347 字 阅读 →
论文解读

Acoustic and Facial Markers of Perceived Conversational Success in Spontaneous Speech

语音情感识别 | 6.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3562 字 阅读 →
论文解读

Acoustic Feedback Cancellation in Hearing Aids Exploiting an Inertial Sensor

音频分类 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4915 字 阅读 →
论文解读

ADH-VA: Adaptive Directed-Hypergraph Convolution with VA Contrastive Learning for Multimodal Conversational Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4727 字 阅读 →
论文解读

Advancing Speech Summarization in Multi-Modal LLMs with Reinforcement Learning

音频问答 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4272 字 阅读 →
论文解读

Affect-Jigsaw: Integrating Core and Peripheral Emotions for Harmonious Fine-Grained Multimodal Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4583 字 阅读 →