论文解读

DBFT-SD: Weakly Supervised Multimodal Detection of Sensitive Audio-Visual Content

音频事件检测 | 8.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 4008 字 阅读 →
论文解读

DDSR-Net: Robust Multimodal Sentiment Analysis via Dynamic Modality Reliability Assessment

语音情感识别 | 6.5/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8029 字 阅读 →
论文解读

Diffemotalk: Audio-Driven Facial Animation with Fine-Grained Emotion Control via Diffusion Models

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5341 字 阅读 →
论文解读

Disentangled Authenticity Representation for Partially Deepfake Audio Localization

音频深度伪造检测 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3989 字 阅读 →
论文解读

DISSR: Disentangling Speech Representation for Degradation-Prior Guided Cross-Domain Speech Restoration

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5058 字 阅读 →
论文解读

DMP-TTS: Disentangled Multi-Modal Prompting for Controllable Text-to-Speech with Chained Guidance

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6161 字 阅读 →
论文解读

Domain-Invariant Representation Learning of Bird Sounds

生物声学 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5091 字 阅读 →
论文解读

DPT-Net: Dual-Path Transformer Network with Hierarchical Fusion for EEG-based Envelope Reconstruction

语音生物标志物 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5126 字 阅读 →
论文解读

DSSR: Decoupling Salient and Subtle Representations Under Missing Modalities for Multimodal Emotion Recognition

情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5035 字 阅读 →
论文解读

Dual Contrastive Learning for Semi-Supervised Domain Adaptation in Bi-Modal Depression Recognition

语音生物标志物 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4365 字 阅读 →
论文解读

Dual Data Scaling for Robust Two-Stage User-Defined Keyword Spotting

语音活动检测 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5107 字 阅读 →
论文解读

Dual-Perspective Multimodal Sentiment Analysis with MoE Fusion: Representation Learning via Semantic Resonance and Divergence

多模态情感分析 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4367 字 阅读 →
论文解读

EchoRAG: A Two-Stage Framework for Audio-Text Retrieval and Temporal Grounding

音频检索 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5031 字 阅读 →
论文解读

Empowering Multimodal Respiratory Sound Classification with Counterfactual Adversarial Debiasing for Out-of-Distribution Robustness

音频分类 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4255 字 阅读 →
论文解读

Face-Voice Association with Inductive Bias for Maximum Class Separation

说话人验证 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4756 字 阅读 →
论文解读

FUSEMOS: Perceptual Evaluation of Text-to-Music Generation with Dual-Encoder Fusion and Ranking-Aware Composite Loss

音乐生成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4701 字 阅读 →
论文解读

GLAP: General Contrastive Audio-Text Pretraining Across Domains and Languages

音频检索 | 8.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4462 字 阅读 →
论文解读

GMS-CAVP: Improving Audio-Video Correspondence with Multi-Scale Constrative and Generative Pretraining

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3524 字 阅读 →
论文解读

Graph-Based Emotion Consensus Perception Learning for Multimodal Emotion Recognition in Conversation

多模态情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4014 字 阅读 →
论文解读

Graph-based Modality Alignment for Robustness in Conversational Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5275 字 阅读 →