论文解读

From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping

扩散模型 | 7.6/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7387 字 阅读 →
论文解读

Group Cognition Learning: Making Everything Better Through Controlled Two-Stage Agents Collaboration

音视频理解 | 7.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6246 字 阅读 →
论文解读

IVQ: Structured and Lightweight Vector Quantization via Binary Hierarchical Composition Inspired by \(\textit{IChing}\)

音频编码 | 8.2/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7794 字 阅读 →
论文解读

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

声源定位 | 8.1/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8152 字 阅读 →
论文解读

Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits

语音识别 | 9.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6415 字 阅读 →
论文解读

LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues

语音交互 | 8.1/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7266 字 阅读 →
论文解读

LightAVSeg: Lightweight Audio-Visual Segmentation

模型压缩 | 6.3/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8803 字 阅读 →
论文解读

MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio

音频理解 | 6.4/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8593 字 阅读 →
论文解读

Multimodal Latent Language Modeling with Next-Token Diffusion

语音合成 | 6.1/10

 · 更新于 2026-09-06 · 约 6 分钟 · 2865 字 阅读 →
论文解读

Multimodal Meta-Verifier with Explicit Structured Recalibration

多模态模型 | 5.2/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6947 字 阅读 →
论文解读

Multiple Choice Learning of Low-Rank Adapters for Language Modeling

多模态模型 | 8/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6764 字 阅读 →
论文解读

Native Active Perception as Reasoning for Omni-Modal Understanding

音视频理解 | 6.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6431 字 阅读 →
论文解读

Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech Separation

音视频语音分离 | 6.2/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9223 字 阅读 →
论文解读

OmniFit: Bridging Modalities via Layer-Adaptive Token Compression for Omnimodal Large Language Models

音视频理解 | 6.3/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8099 字 阅读 →
论文解读

PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios

音视频问答 | 7.3/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7522 字 阅读 →
论文解读

PRIM:Cooperative Dynamic Token Compression for Efficient Large Multimodal Models

音视频理解 | 3.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5669 字 阅读 →
论文解读

Probing Cross-modal Information Hubs in Audio-Visual LLMs

音视频理解 | 7.2/10

 · 更新于 2026-09-06 · 约 6 分钟 · 2684 字 阅读 →
论文解读

SAM Audio: Segment Anything in Audio

音频分离 | 9.2/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6829 字 阅读 →
论文解读

Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis

音视频生成 | 7.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6404 字 阅读 →
论文解读

SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering

音频修复 | 8/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8353 字 阅读 →