论文解读

MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue

跨模态 | 6.5/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6048 字 阅读 →
论文解读

VISAFF: Speaker-Centered Visual Affective Feature Learning for Emotion Recognition in Conversation

对话情感识别 | 7.4/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6752 字 阅读 →
论文解读

Caption and Audio-Guided Video Representation Learning with Gated Attention for Partially Relevant Video Retrieval

视频检索 | 7.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4469 字 阅读 →