论文解读

CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks

音视频理解 | 8.3/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8159 字 阅读 →
论文解读

ConsMSA: Semantic Distribution Consistency Learning for Multimodal Sentiment Analysis

多模态模型 | 6.1/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5246 字 阅读 →
论文解读

DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation

音视频生成 | 8/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9827 字 阅读 →
论文解读

E-VAds: An E-commerce Short Videos Understanding Benchmark for MLLMs

音视频问答 | 6.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8438 字 阅读 →
论文解读

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

音视频理解 | 6.3/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6514 字 阅读 →
论文解读

Efficient Distributed MLLM Training with Cornstarch

音视频理解 | 7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5041 字 阅读 →
论文解读

FakeWorld 1.0: An Omni-modal Benchmark for Fake Media and Content

可解释性 | 6.1/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4956 字 阅读 →
论文解读

From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping

扩散模型 | 7.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7387 字 阅读 →
论文解读

Group Cognition Learning: Making Everything Better Through Controlled Two-Stage Agents Collaboration

音视频理解 | 7.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6246 字 阅读 →
论文解读

IVQ: Structured and Lightweight Vector Quantization via Binary Hierarchical Composition Inspired by \(\textit{IChing}\)

音频编码 | 8.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7794 字 阅读 →
论文解读

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

声源定位 | 8.1/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8152 字 阅读 →
论文解读

Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits

语音识别 | 9.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6415 字 阅读 →
论文解读

LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues

语音交互 | 8.1/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7266 字 阅读 →
论文解读

LightAVSeg: Lightweight Audio-Visual Segmentation

模型压缩 | 6.3/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8803 字 阅读 →
论文解读

MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio

音频理解 | 6.4/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8593 字 阅读 →
论文解读

Multimodal Latent Language Modeling with Next-Token Diffusion

语音合成 | 6.1/10

 · 更新于 2026-09-25 · 约 6 分钟 · 2865 字 阅读 →
论文解读

Multimodal Meta-Verifier with Explicit Structured Recalibration

多模态模型 | 5.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6947 字 阅读 →
论文解读

Multiple Choice Learning of Low-Rank Adapters for Language Modeling

多模态模型 | 8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6764 字 阅读 →
论文解读

Native Active Perception as Reasoning for Omni-Modal Understanding

音视频理解 | 6.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6431 字 阅读 →
论文解读

Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech Separation

音视频语音分离 | 6.2/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9223 字 阅读 →