论文解读

Rethinking Entity Disambiguation in Complex Modalities

实体消歧 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4884 字 阅读 →
论文解读

Rethinking Music Captioning with Music Metadata LLMS

音乐理解 | 7.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5562 字 阅读 →
论文解读

Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models

语音情感识别 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4289 字 阅读 →
论文解读

Selective Hub Fusion with Modality-Heterogeneous Experts for Multimodal Emotion Recognition

多模态模型 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4869 字 阅读 →
论文解读

Semantic-Guided Pseudo-Feature Attention Network for Audio-Visual Zero-Shot Learning

音频分类 零样本学习 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4921 字 阅读 →
论文解读

Session-Level Spoken Language Assessment with A Multimodal Foundation Model Via Multi-Target Learning

语音评估 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4959 字 阅读 →
论文解读

SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

音频问答 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5286 字 阅读 →
论文解读

SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training

音频检索 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4546 字 阅读 →
论文解读

Sparse-View Visual-Acoustic Latent Learning for Novel-View Audio Synthesis

空间音频 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5274 字 阅读 →
论文解读

SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis

医疗AI | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4415 字 阅读 →
论文解读

Spiking Temporal-Enhanced Network for Zero-Shot Audio-Visual Learning

音频分类 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4341 字 阅读 →
论文解读

ST-HNTM: Joint Speech-Text Neural Topic Modeling on the Hypersphere

主题建模 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4767 字 阅读 →
论文解读

Staged Diffusion with Hybrid Mixture-of-Experts (MOE) for Multimodal Sentiment Analysis

语音情感识别 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4137 字 阅读 →
论文解读

Still Thinking or Stopped Talking? Dialogue Silence Intention Classification Using Multimodal Large Language Model

语音对话系统 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3789 字 阅读 →
论文解读

Streamingbench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

基准测试 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3905 字 阅读 →
论文解读

SURE: Synergistic Uncertainty-Aware Reasoning for Multimodal Emotion Recognition in Conversations

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4997 字 阅读 →
论文解读

SynaSpot: A Lightweight, Streaming Multi-modal Framework for Keyword Spotting with Audio-Text Synergy

关键词检测 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3921 字 阅读 →
论文解读

Temporal-Spatial Decouple Before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis

情感分析 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6191 字 阅读 →
论文解读

The Curious Case of Visual Grounding: Different Effects for Speech-and Text-Based Language Encoders

模型评估 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4508 字 阅读 →
论文解读

The Synergistic Role of Audio and Large Video-Language Model in Source-Free Video Domain Adaptation

领域适应 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4434 字 阅读 →