论文解读

From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection

From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs

FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

Group Cognition Learning: Making Everything Better Through Controlled Two-Stage Agents Collaboration

Group Cognition Learning: Making Everything Better Through Controlled Two-Stage Agents Collaboration

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

Hearing Without Noticing? Attention-Aware Stealthy Black-box Adversarial Audio Attacks

Hearing Without Noticing? Attention-Aware Stealthy Black-box Adversarial Audio Attacks

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →
论文解读

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

HyperPotter: Spell the Charm of High-Order Interactions in Audio Deepfake Detection

HyperPotter: Spell the Charm of High-Order Interactions in Audio Deepfake Detection

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

INFER: Learning Implicit Neural Frequency Response Fields for Confined Acoustic Environments

INFER: Learning Implicit Neural Frequency Response Fields for Confined Acoustic Environments

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

IVQ: Structured and Lightweight Vector Quantization via Binary Hierarchical Composition Inspired by IChing

IVQ: Structured and Lightweight Vector Quantization via Binary Hierarchical Composition Inspired by IChing

 · 更新于 2026-09-09 · 约 1 分钟 · 39 字 阅读 →
论文解读

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits

Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits

 · 更新于 2026-09-09 · 约 1 分钟 · 38 字 阅读 →
论文解读

LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues

LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues

 · 更新于 2026-09-09 · 约 1 分钟 · 38 字 阅读 →
论文解读

Language Model Augmented Semi-Supervised Statistical Inference

Language Model Augmented Semi-Supervised Statistical Inference

 · 更新于 2026-09-09 · 约 1 分钟 · 32 字 阅读 →
论文解读

Learning Tight Rejection Boundaries without Negatives for Strict One-Class Audio Deepfake Detection

Learning Tight Rejection Boundaries without Negatives for Strict One-Class Audio Deepfake Detection

 · 更新于 2026-09-09 · 约 1 分钟 · 38 字 阅读 →
论文解读

LightAVSeg: Lightweight Audio-Visual Segmentation

LightAVSeg: Lightweight Audio-Visual Segmentation

 · 更新于 2026-09-09 · 约 1 分钟 · 30 字 阅读 →
论文解读

Listening Through the Noise: Cauchy-Driven Diffusion Bridges for Robust Gastrointestinal Auscultation and Clinical Benchmarking

Listening Through the Noise: Cauchy-Driven Diffusion Bridges for Robust Gastrointestinal Auscultation and Clinical Benchmarking

 · 更新于 2026-09-09 · 约 1 分钟 · 40 字 阅读 →
论文解读

Long Grounded Thoughts: Synthesizing Grounded Visual Problems and Distilling Reasoning Chains at Scale

Long Grounded Thoughts: Synthesizing Grounded Visual Problems and Distilling Reasoning Chains at Scale

 · 更新于 2026-09-09 · 约 1 分钟 · 39 字 阅读 →
论文解读

LynX: Token Interface Alignment for Video+X LLMs

** | 7.5/10

 · 更新于 2026-09-09 · 约 1 分钟 · 63 字 阅读 →
论文解读

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio

MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio

 · 更新于 2026-09-09 · 约 1 分钟 · 36 字 阅读 →
论文解读

MetaBio: Learning from metadata for bioacoustics foundation models

MetaBio: Learning from metadata for bioacoustics foundation models

 · 更新于 2026-09-09 · 约 1 分钟 · 34 字 阅读 →