Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio
Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio
Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio
HyperPotter: Spell the Charm of High-Order Interactions in Audio Deepfake Detection
INFER: Learning Implicit Neural Frequency Response Fields for Confined Acoustic Environments
IVQ: Structured and Lightweight Vector Quantization via Binary Hierarchical Composition Inspired by IChing
JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments
Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits
LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
Language Model Augmented Semi-Supervised Statistical Inference
Learning Tight Rejection Boundaries without Negatives for Strict One-Class Audio Deepfake Detection
LightAVSeg: Lightweight Audio-Visual Segmentation
Listening Through the Noise: Cauchy-Driven Diffusion Bridges for Robust Gastrointestinal Auscultation and Clinical Benchmarking
Long Grounded Thoughts: Synthesizing Grounded Visual Problems and Distilling Reasoning Chains at Scale
** | 7.5/10
MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio
MetaBio: Learning from metadata for bioacoustics foundation models
MFCL Audio: An Audio Function Calling Evaluation for Large Language Models
MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts
视频行为识别 | 5.9/10