From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
Group Cognition Learning: Making Everything Better Through Controlled Two-Stage Agents Collaboration
Hearing Without Noticing? Attention-Aware Stealthy Black-box Adversarial Audio Attacks
Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio
HyperPotter: Spell the Charm of High-Order Interactions in Audio Deepfake Detection
INFER: Learning Implicit Neural Frequency Response Fields for Confined Acoustic Environments
IVQ: Structured and Lightweight Vector Quantization via Binary Hierarchical Composition Inspired by IChing
JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments
Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits
LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
Language Model Augmented Semi-Supervised Statistical Inference
Learning Tight Rejection Boundaries without Negatives for Strict One-Class Audio Deepfake Detection
LightAVSeg: Lightweight Audio-Visual Segmentation
Listening Through the Noise: Cauchy-Driven Diffusion Bridges for Robust Gastrointestinal Auscultation and Clinical Benchmarking
Long Grounded Thoughts: Synthesizing Grounded Visual Problems and Distilling Reasoning Chains at Scale
** | 7.5/10
MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio
MetaBio: Learning from metadata for bioacoustics foundation models