论文解读

Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs

模型评估 | 8.0/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10630 字 阅读 →
论文解读

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

多模态模型评估 | 5.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9061 字 阅读 →
论文解读

jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Composition

多模态检索 | 7.5/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8909 字 阅读 →
论文解读

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation

音视频生成 | 6.9/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8863 字 阅读 →
论文解读

AllocMV: Optimal Resource Allocation for Music Video Generation via Structured Persistent State

音乐视频生成 | 4.8/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8053 字 阅读 →
论文解读

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing

音频事件检测 | 5.8/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9380 字 阅读 →
论文解读

Mitigating Multimodal Inconsistency via Cognitive Dual-Pathway Reasoning for Intent Recognition

意图识别 | 7.0/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9463 字 阅读 →
论文解读

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

基准测试 | 6.5/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8876 字 阅读 →
论文解读

Probing Cross-modal Information Hubs in Audio-Visual LLMs

模型分析 | 6.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9363 字 阅读 →
论文解读

Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought

音视频问答 | 6.0/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9627 字 阅读 →
论文解读

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes

MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes

 · 更新于 2026-09-25 · 约 14 分钟 · 7004 字 阅读 →
论文解读

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs

 · 更新于 2026-09-25 · 约 18 分钟 · 8609 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-11

共分析 12 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 41 分钟 · 20158 字 阅读 →
论文解读

Audio-Visual Intelligence in Large Foundation Models

跨模态 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4580 字 阅读 →
论文解读

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

移动代理 | 6.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5789 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-09

共分析 3 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 10 分钟 · 4857 字 阅读 →
论文解读

Modality-Aware Contrastive and Uncertainty-Regularized Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7828 字 阅读 →
论文解读

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions

音频质量评估 | 8.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6589 字 阅读 →
论文解读

OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models

音频分类 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4468 字 阅读 →
论文解读

To Fuse or to Drop? Dual-Path Learning for Resolving Modality Conflicts in Multimodal Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6769 字 阅读 →