论文解读

SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations

音频理解 | 8.4/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7325 字 阅读 →
论文解读

Speech-Audio Compositional Attacks on Multimodal LLMs and Their Defense with SALMONN-Guard

音频理解 | 8.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6409 字 阅读 →
论文解读

Spherical Procrustes Alignment for Reliable Medical Audio Diagnosis

音频分类 | 8.2/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8092 字 阅读 →
论文解读

Stable Spectral Copula Alignment for Robust Multimodal Learning

鲁棒性 | 5.2/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7374 字 阅读 →
论文解读

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

音频生成 | 7.9/10

 · 更新于 2026-09-06 · 约 20 分钟 · 9942 字 阅读 →
论文解读

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

音视频生成 | 6.8/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7540 字 阅读 →
论文解读

Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage

流式处理 | 7.2/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9066 字 阅读 →
论文解读

SURF: Separation via Unsupervised Remixing Flow

语音分离 | 6.2/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8040 字 阅读 →
论文解读

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

音视频生成 | 7.9/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8106 字 阅读 →
论文解读

TextME: Bridging Unseen Modalities Through Text Descriptions

TextME: Bridging Unseen Modalities Through Text Descriptions

 · 更新于 2026-09-06 · 约 16 分钟 · 7828 字 阅读 →
论文解读

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

知识蒸馏 | 7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6068 字 阅读 →
论文解读

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

音视频理解 | 9.4/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7829 字 阅读 →
论文解读

TMD-Bench: A Multi-Level Evaluation Paradigm for Music–Dance Co-Generation

音视频生成 | 7.7/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5578 字 阅读 →
论文解读

Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer

音视频生成 | 6.6/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6714 字 阅读 →
论文解读

Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition

音视频理解 | 5.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6064 字 阅读 →
论文解读

Two-dimensional quantization for geometry-aware audio coding

语音编码 | 7.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5851 字 阅读 →
论文解读

UltraLIF: Fully Differentiable Spiking Neural Networks via Ultradiscretization and Max-Plus Algebra

音频分类 | 6.7/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7315 字 阅读 →
论文解读

UniFLoW: Universal Multi-Modal Federated LoRA Fine-Tuning Framework with Analytical Aggregation

音视频问答 | 4.4/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8116 字 阅读 →
论文解读

Universal Algorithm-Implicit Learning

音频分类 | 6.5/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7593 字 阅读 →
论文解读

Unlocking Cross-Modal Biosignal Synthesis: A Temporally-Aware VAE-Diffusion Model

Unlocking Cross-Modal Biosignal Synthesis: A Temporally-Aware VAE-Diffusion Model

 · 更新于 2026-09-06 · 约 24 分钟 · 11651 字 阅读 →