论文解读

Scaling Laws in Model Fine-tuning for Audio DeepFake Detection

Scaling Laws in Model Fine-tuning for Audio DeepFake Detection

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →
论文解读

Scaling Transformers for End-to-End Discrete Audio Tokenization

Scaling Transformers for End-to-End Discrete Audio Tokenization

 · 更新于 2026-09-09 · 约 1 分钟 · 33 字 阅读 →
论文解读

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment

 · 更新于 2026-09-09 · 约 1 分钟 · 34 字 阅读 →
论文解读

Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis

Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis

 · 更新于 2026-09-09 · 约 1 分钟 · 33 字 阅读 →
论文解读

Simultaneous Speech-to-Speech Translation Without Aligned Data

Simultaneous Speech-to-Speech Translation Without Aligned Data

 · 更新于 2026-09-09 · 约 1 分钟 · 32 字 阅读 →
论文解读

SONAR: Spectral‑Contrastive Audio Residuals for Generalizable Deepfake Detection

SONAR: Spectral‑Contrastive Audio Residuals for Generalizable Deepfake Detection

 · 更新于 2026-09-09 · 约 1 分钟 · 53 字 阅读 →
论文解读

SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering

SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering

 · 更新于 2026-09-09 · 约 1 分钟 · 34 字 阅读 →
论文解读

Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech

Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech

 · 更新于 2026-09-09 · 约 1 分钟 · 34 字 阅读 →
论文解读

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations

SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

Speech-Audio Compositional Attacks on Multimodal LLMs and Their Defense with SALMONN-Guard

Speech-Audio Compositional Attacks on Multimodal LLMs and Their Defense with SALMONN-Guard

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

Spherical Procrustes Alignment for Reliable Medical Audio Diagnosis

Spherical Procrustes Alignment for Reliable Medical Audio Diagnosis

 · 更新于 2026-09-09 · 约 1 分钟 · 34 字 阅读 →
论文解读

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

 · 更新于 2026-09-09 · 约 1 分钟 · 35 字 阅读 →
论文解读

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

 · 更新于 2026-09-09 · 约 1 分钟 · 37 字 阅读 →
论文解读

Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage

Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage

 · 更新于 2026-09-09 · 约 1 分钟 · 38 字 阅读 →
论文解读

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

 · 更新于 2026-09-09 · 约 1 分钟 · 33 字 阅读 →
论文解读

tau-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains

tau-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains

 · 更新于 2026-09-09 · 约 0 分钟 · 0 字 阅读 →
论文解读

TextME: Bridging Unseen Modalities Through Text Descriptions

TextME: Bridging Unseen Modalities Through Text Descriptions

 · 更新于 2026-09-09 · 约 1 分钟 · 33 字 阅读 →
论文解读

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

 · 更新于 2026-09-09 · 约 1 分钟 · 40 字 阅读 →
论文解读

TMD-Bench: A Multi-Level Evaluation Paradigm for Music–Dance Co-Generation

TMD-Bench: A Multi-Level Evaluation Paradigm for Music–Dance Co-Generation

 · 更新于 2026-09-09 · 约 1 分钟 · 44 字 阅读 →