论文解读

Staged Diffusion with Hybrid Mixture-of-Experts (MOE) for Multimodal Sentiment Analysis

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4137 字 阅读 →
论文解读

Still Thinking or Stopped Talking? Dialogue Silence Intention Classification Using Multimodal Large Language Model

语音对话系统 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3789 字 阅读 →
论文解读

Streamingbench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

基准测试 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3905 字 阅读 →
论文解读

SURE: Synergistic Uncertainty-Aware Reasoning for Multimodal Emotion Recognition in Conversations

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4997 字 阅读 →
论文解读

SynaSpot: A Lightweight, Streaming Multi-modal Framework for Keyword Spotting with Audio-Text Synergy

关键词检测 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3921 字 阅读 →
论文解读

Temporal-Spatial Decouple Before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis

情感分析 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6191 字 阅读 →
论文解读

The Curious Case of Visual Grounding: Different Effects for Speech-and Text-Based Language Encoders

模型评估 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4508 字 阅读 →
论文解读

The Synergistic Role of Audio and Large Video-Language Model in Source-Free Video Domain Adaptation

领域适应 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4434 字 阅读 →
论文解读

TinyMU: A Compact Audio-Language Model for Music Understanding

音乐理解 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5011 字 阅读 →
论文解读

Towards Effective Negation Modeling in Joint Audio-Text Models for Music

音乐理解 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4421 字 阅读 →
论文解读

Towards Multi-View Hierarchical Video-to-Piano Generation with MIDI Guidance

音乐生成 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4459 字 阅读 →
论文解读

Tpeformer: Temporal Patch Embedding Transformer

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4746 字 阅读 →
论文解读

Training-Free Multimodal Guidance for Video to Audio Generation

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4470 字 阅读 →
论文解读

Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior Distillation

音视频 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4803 字 阅读 →
论文解读

UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4909 字 阅读 →
论文解读

UVT-LM: Unifying Visual and Tactile Perception with Language Model

跨模态 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4823 字 阅读 →
论文解读

VMSP: Video-to-Music Generation with Two-Stage Alignment and Synthesis

音乐生成 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4643 字 阅读 →
论文解读

VT-Heads: Voice Cloning and Talking Head Generation from Text Based on V-DiT

视频生成 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3563 字 阅读 →
论文解读

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

音频问答 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3527 字 阅读 →
论文解读

When Audio Matters: A Lightweight, Hierarchical Fusion Model for Speech and Non-Verbal Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4926 字 阅读 →