论文解读

CustomDancer: Customized Dance Recommendation by Text-Dance Retrieval

音频检索 音乐理解 | 6.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5689 字 阅读 →
论文解读

MARS-Sep: Multimodal-Aligned Reinforced Sound Separation

语音分离 | 7.5/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12209 字 阅读 →
论文解读

MindMix: A Multimodal Foundation Model for Auditory Perception Decoding via Deep Neural-Acoustic Alignment

音频分类 | 9.0/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9208 字 阅读 →
论文解读

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

多模态模型 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5484 字 阅读 →
论文解读

RoboKA: KAN Informed Multimodal Learning for RoboCall Surveillance System

语音伪造检测 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4899 字 阅读 →
论文解读

SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization

音频检索 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4249 字 阅读 →
论文解读

The Deleuzian Representation Hypothesis

模型可解释性 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4780 字 阅读 →
论文解读

WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

音频检索 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4709 字 阅读 →
论文解读

Beyond Instance-Level Alignment: Dual-Level Optimal Transport for Audio-Text Retrieval

音频检索 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4293 字 阅读 →
论文解读

Learning multimodal dictionary decompositions with group-sparse autoencoders

跨模态 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4031 字 阅读 →
论文解读

MindMix: A Multimodal Foundation Model for Auditory Perception Decoding via Deep Neural-Acoustic Alignment

音频检索 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5061 字 阅读 →
论文解读

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM

音频问答 | 8.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6263 字 阅读 →
论文解读

SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization

音频检索 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4253 字 阅读 →
论文解读

Unified Multi-Modal Interactive and Reactive 3D Motion Generation via Rectified Flow

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4970 字 阅读 →
论文解读

WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM

音频检索 视频检索 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5230 字 阅读 →
论文解读

Diffusion Reconstruction towards Generalizable Audio Deepfake Detection

音频深度伪造检测 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4407 字 阅读 →
论文解读

Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection

音频深度伪造检测 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4571 字 阅读 →
论文解读

A Hybrid Convolution-Mamba Network with Tone-Octave Contrastive Learning for Stratified Semi-Supervised Singing Melody Extraction

歌唱旋律提取 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6269 字 阅读 →
论文解读

A LLM-Driven Acoustic Semantic Enriched Framework for Underwater Acoustic Target Recognition

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4601 字 阅读 →
论文解读

A Metric Learning Approach to Heart Murmur Detection from Phonocardiogram Recordings

音频分类 | 7.7/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3727 字 阅读 →