论文解读

Investigating Modality Contribution in Audio LLMs for Music

模型评估 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3688 字 阅读 →
论文解读

Learning What to Hear: Boosting Sound-Source Association for Robust Audiovisual Instance Segmentation

音视频实例分割 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4721 字 阅读 →
论文解读

LETPAV: Lexicon-Enhanced Text with Progressive Audio-Visual Fusion for Multimodal Sentiment Analysis

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4653 字 阅读 →
论文解读

Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models

语音识别 | 6.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4691 字 阅读 →
论文解读

Leveraging Large Multimodal Models for Audio-Video Deepfake Detection: A Pilot Study

音频深度伪造检测 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4328 字 阅读 →
论文解读

Leveraging prediction entropy for Automatic prompt weighting in Zero-Shot Audio-Language Classification

音频分类 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4393 字 阅读 →
论文解读

MAG: Multi-Modal Aligned Autoregressive Co-Speech Gesture Generation Without Vector Quantization

音频生成 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4829 字 阅读 →
论文解读

MCF: Text LLMS for Multimodal Emotional Causality

情感分析 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4895 字 阅读 →
论文解读

MECap-R1: Emotion-Aware Policy with Reinforcement Learning for Multimodal Emotion Captioning

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5068 字 阅读 →
论文解读

MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding

音乐理解 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4303 字 阅读 →
论文解读

Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4158 字 阅读 →
论文解读

Mitigating Language Prior-Induced Hallucinations via Bi-Level Contrastive Decoding

多模态模型 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3934 字 阅读 →
论文解读

Mitigating Shared-Private Branch Imbalance via Dual-Branch Rebalancing for Multimodal Sentiment Analysis

多模态模型 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5421 字 阅读 →
论文解读

Mixture of Experts for Recognizing Depression from Interview and Reading Tasks

语音生物标志物 | 6.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5201 字 阅读 →
论文解读

ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations

语音情感识别 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4331 字 阅读 →
论文解读

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

语音分离 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4749 字 阅读 →
论文解读

MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models

基准测试 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4488 字 阅读 →
论文解读

Modeling Both Intra- And Inter-Utterance Variability for Conversational Emotion Recognition

语音情感识别 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4447 字 阅读 →
论文解读

MSANET: Multi-Scale Semantic Aggregation Network for Brain-Assisted Speech Enhancement in Multi-Speaker Conditions

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4444 字 阅读 →
论文解读

MSCT: Differential Cross-Modal Attention for Deepfake Detection

音频深度伪造检测 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3968 字 阅读 →