论文解读

Inter-Dialog Contrastive Learning for Multimodal Emotion Recognition in Conversations

语音情感识别 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4414 字 阅读 →
论文解读

KSDIFF: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation

音频生成 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4327 字 阅读 →
论文解读

LETPAV: Lexicon-Enhanced Text with Progressive Audio-Visual Fusion for Multimodal Sentiment Analysis

语音情感识别 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4653 字 阅读 →
论文解读

Leveraging Large Multimodal Models for Audio-Video Deepfake Detection: A Pilot Study

音频深度伪造检测 | 7.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4328 字 阅读 →
论文解读

Look, Listen and Segment: Towards Weakly Supervised Audio-Visual Semantic Segmentation

音视频 | 7.0/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4187 字 阅读 →
论文解读

MCI-OTFusion: A Multimodal Model for MCI Detection and Cognitive Score Prediction

轻度认知障碍检测 | 6.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5072 字 阅读 →
论文解读

Mitigating Shared-Private Branch Imbalance via Dual-Branch Rebalancing for Multimodal Sentiment Analysis

多模态模型 | 7.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5421 字 阅读 →
论文解读

MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models

基准测试 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4488 字 阅读 →
论文解读

Motionbeat: Motion-Aligned Music Representation via Embodied Contrastive Learning and Bar-Equivariant Contact-Aware Encoding

舞蹈生成 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4092 字 阅读 →
论文解读

Multi-Scale Physiologically-Motivated Alignment for Auditory Attention Decoding

听觉注意力解码 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4347 字 阅读 →
论文解读

Multimodal Fusion-Based IPCLIP Network for Mixed Reality Surgical Assistance

多模态模型 | 6.5/10

 · 更新于 2026-09-24 · 约 7 分钟 · 3263 字 阅读 →
论文解读

Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4714 字 阅读 →
论文解读

Natural Language to Spatial Audio Parameters: Lightweight Deterministic Rendering for Creative Authoring

空间音频 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4737 字 阅读 →
论文解读

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

多模态模型 | 8.5/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6861 字 阅读 →
论文解读

NeuroSIFT: A Biologically-Inspired Framework with Explicit Signal-Noise Separation for Robust Multimodal Emotion Recognition

多模态情感识别 | 8.0/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3891 字 阅读 →
论文解读

RCAL: Reinforced Cross-Modal Alignment for Multimodal Sentiment Analysis with Sparse Visual Frames

多模态模型 | 8.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4562 字 阅读 →
论文解读

Reliable AI via Age-Balanced Validation: Fair Model Selection for Parkinson’s Detection from Voice

语音生物标志物 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4513 字 阅读 →
论文解读

Savgbench: Benchmarking Spatially Aligned Audio-Video Generation

基准测试 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4030 字 阅读 →
论文解读

Selective Hub Fusion with Modality-Heterogeneous Experts for Multimodal Emotion Recognition

多模态模型 | 6.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4869 字 阅读 →
论文解读

Sounds that Shape: Audio-Driven 3D Mesh Generation with Attribute-Decoupled Score Distillation Sampling

音频生成 | 7.0/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3847 字 阅读 →