论文解读

HarmoNet: Music Grounding by Short Video via Harmonic Resample and Dynamic Sparse Alignment

音乐检索 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4256 字 阅读 →
论文解读

HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-Based TTS

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4633 字 阅读 →
论文解读

Improving Binaural Distance Estimation in Reverberant Rooms Through Contrastive And Multi-Task Learning

声源定位 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4718 字 阅读 →
论文解读

Inter-Dialog Contrastive Learning for Multimodal Emotion Recognition in Conversations

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4414 字 阅读 →
论文解读

Learning Domain-Robust Bioacoustic Representations for Mosquito Species Classification with Contrastive Learning and Distribution Alignment

生物声学 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5151 字 阅读 →
论文解读

LETPAV: Lexicon-Enhanced Text with Progressive Audio-Visual Fusion for Multimodal Sentiment Analysis

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4653 字 阅读 →
论文解读

Leveraging Whisper Embeddings For Audio-Based Lyrics Matching

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4853 字 阅读 →
论文解读

Lightweight and Generalizable Acoustic Scene Representations Via Contrastive Fine-Tuning and Distillation

音频场景理解 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5226 字 阅读 →
论文解读

Look, Listen and Segment: Towards Weakly Supervised Audio-Visual Semantic Segmentation

音视频 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4187 字 阅读 →
论文解读

MAG: Multi-Modal Aligned Autoregressive Co-Speech Gesture Generation Without Vector Quantization

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4829 字 阅读 →
论文解读

Malefa: Multi-Granularity Learning and Effective False Alarm Suppression for Zero-Shot Keyword Spotting

零样本关键词检测 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4936 字 阅读 →
论文解读

MC-MRX: Reference- and Midi-Guided Music Source Extraction with Contrastive Learning

音乐源提取 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5478 字 阅读 →
论文解读

Mitigating Language Prior-Induced Hallucinations via Bi-Level Contrastive Decoding

多模态模型 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3934 字 阅读 →
论文解读

Mitigating Shared-Private Branch Imbalance via Dual-Branch Rebalancing for Multimodal Sentiment Analysis

多模态模型 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5421 字 阅读 →
论文解读

Motionbeat: Motion-Aligned Music Representation via Embodied Contrastive Learning and Bar-Equivariant Contact-Aware Encoding

舞蹈生成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4092 字 阅读 →
论文解读

Multi-Scale Physiologically-Motivated Alignment for Auditory Attention Decoding

听觉注意力解码 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4347 字 阅读 →
论文解读

Noise-Robust Contrastive Learning with an MFCC-Conformer for Coronary Artery Disease Detection

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5224 字 阅读 →
论文解读

PADAM: Perceptual Audio Defect Assessment Model

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4814 字 阅读 →
论文解读

Prototype-Guided Cross-Modal Contrastive Learning for Continual Audio-Visual Sound Separation

语音分离 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4252 字 阅读 →
论文解读

Rationale-Guided Learning for Multimodal Emotion Recognition

语音情感识别 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4768 字 阅读 →