论文解读

Audio-Image Cross-Modal Retrieval with Onomatopoeic Images

音频检索 | 7/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6984 字 阅读 →
论文解读

Physics-Based iOCT Sonification for Real-time Interaction Awareness in Subretinal Injection

医疗音频 | 6.5/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7411 字 阅读 →
论文解读

Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs

音视频 | 7.0/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8769 字 阅读 →
论文解读

Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs

模型评估 | 8.0/10

 · 更新于 2026-09-24 · 约 22 分钟 · 10630 字 阅读 →
论文解读

CORTEG: Foundation Models Enable Cross-Modality Representation Transfer from Scalp to Intracranial Brain Recordings

脑机接口 | 6.5/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9166 字 阅读 →
论文解读

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing

音频事件检测 | 5.8/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9380 字 阅读 →
论文解读

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries

音频检索 | 6.0/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9517 字 阅读 →
论文解读

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

基准测试 | 6.0/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7890 字 阅读 →
论文解读

Anisotropic Modality Align

Anisotropic Modality Align

 · 更新于 2026-09-24 · 约 16 分钟 · 7980 字 阅读 →
论文解读

Do Joint Audio-Video Generation Models Understand Physics?

Do Joint Audio-Video Generation Models Understand Physics?

 · 更新于 2026-09-24 · 约 18 分钟 · 8531 字 阅读 →
论文解读

Audio-Visual Intelligence in Large Foundation Models

跨模态 | 8.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4580 字 阅读 →
论文解读

MG-Former: A Transformer-Based Framework for Music-Driven 3D Conducting Gesture Generation

音乐生成 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4252 字 阅读 →
论文解读

Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time

多模态幻觉缓解 | 7.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5400 字 阅读 →
论文解读

A cross-species neural foundation model for end-to-end speech decoding

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5724 字 阅读 →
论文解读

Beyond Instance-Level Alignment: Dual-Level Optimal Transport for Audio-Text Retrieval

音频检索 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4524 字 阅读 →
论文解读

Closing the Gap Between Text and Speech Understanding in LLMs

语音大模型 | 8.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4518 字 阅读 →
论文解读

Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning

多模态推理 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4438 字 阅读 →
论文解读

Group Cognition Learning: Making Everything Better Through Governed Two-Stage Agents Collaboration

多模态模型 | 7.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5604 字 阅读 →
论文解读

MARS-Sep: Multimodal-Aligned Reinforced Sound Separation

语音分离 | 7.5/10

 · 更新于 2026-09-24 · 约 25 分钟 · 12209 字 阅读 →
论文解读

MindMix: A Multimodal Foundation Model for Auditory Perception Decoding via Deep Neural-Acoustic Alignment

音频分类 | 9.0/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9208 字 阅读 →