论文解读

A Unsupervised Domain Adaptation Framework For Semi-Supervised Melody Extraction Using Confidence Matrix Replace and Nearest Neighbour Supervision

音乐信息检索 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4391 字 阅读 →
论文解读

ACIR-MACL: Effective Multimodal Sentiment Analysis via Attention-Based Causal Intervention Regularization and Multi-Aspect Contrastive Learning

情感分析 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5347 字 阅读 →
论文解读

Adaptive Embedding Fusion with Contrastive Learning for Robust Fully Few-Shot Class-Incremental Audio Classification

音频分类 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5041 字 阅读 →
论文解读

ADH-VA: Adaptive Directed-Hypergraph Convolution with VA Contrastive Learning for Multimodal Conversational Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4727 字 阅读 →
论文解读

ALMA-Chor: Leveraging Audio-Lyric Alignment with Mamba for Chorus Detection

音乐信息检索 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4379 字 阅读 →
论文解读

An Anomaly-Aware and Audio-Enhanced Dual-Pathway Framework for Alzheimer’s Disease Progression Classification

语音生物标志物 | 7.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5640 字 阅读 →
论文解读

AnimalCLAP: Taxonomy-Aware Language-Audio Pretraining for Species Recognition and Trait Inference

音频分类 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4960 字 阅读 →
论文解读

ATOM: Adaptive Token-Level Optimal Transport Mixup for Speech Translation

语音翻译 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4810 字 阅读 →
论文解读

Audio-Guided Multimodal Approach for Fine-Grained Alignment and Boundary Modeling in Active Speaker Detection

说话人检测 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4161 字 阅读 →
论文解读

Audio-Visual Deepfake Generation and Detection: An Exploratory Survey

音频深度伪造检测 | 6.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3233 字 阅读 →
论文解读

AUDIOCARDS: Structured Metadata Improves Audio Language Models for Sound Design

音频检索 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5242 字 阅读 →
论文解读

Automatic Music Sample Identification with Multi-Track Contrastive Learning

音频检索 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4382 字 阅读 →
论文解读

BEST-STD 2.0: Balanced and Efficient Speech Tokenizer for Spoken Term Detection

音频检索 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5339 字 阅读 →
论文解读

Bridging the Semantic Gap: Cross-Attentive Fusion for Joint Acoustic-Semantic Speech Quality Assessment

语音质量评估 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5819 字 阅读 →
论文解读

Caption and Audio-Guided Video Representation Learning with Gated Attention for Partially Relevant Video Retrieval

视频检索 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4469 字 阅读 →
论文解读

Contrastive Timbre Representations for Musical Instrument And Synthesizer Retrieval

音频检索 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3791 字 阅读 →
论文解读

Controllable Embedding Transformation for Mood-Guided Music Retrieval

音乐检索 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4452 字 阅读 →
论文解读

CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content

跨模态检索 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4170 字 阅读 →
论文解读

Cross-Domain Contrastive Learning with Dynamic Threshold Calibration for Source Speaker Tracing

说话人验证 | 8.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3746 字 阅读 →
论文解读

Curriculum Learning with Contrastive Loss for Lightweight Speaker Verification

说话人验证 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4588 字 阅读 →