Scaling Laws in Model Fine-tuning for Audio DeepFake Detection
Scaling Laws in Model Fine-tuning for Audio DeepFake Detection
Scaling Laws in Model Fine-tuning for Audio DeepFake Detection
Scaling Transformers for End-to-End Discrete Audio Tokenization
Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment
Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
Simultaneous Speech-to-Speech Translation Without Aligned Data
SONAR: Spectral‑Contrastive Audio Residuals for Generalizable Deepfake Detection
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
Speech-Audio Compositional Attacks on Multimodal LLMs and Their Defense with SALMONN-Guard
Spherical Procrustes Alignment for Reliable Medical Audio Diagnosis
STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage
T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation
tau-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
TextME: Bridging Unseen Modalities Through Text Descriptions
The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning
TMD-Bench: A Multi-Level Evaluation Paradigm for Music–Dance Co-Generation