论文解读

Mitigating Shared-Private Branch Imbalance via Dual-Branch Rebalancing for Multimodal Sentiment Analysis

多模态模型 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5421 字 阅读 →
论文解读

Mixture of Experts for Recognizing Depression from Interview and Reading Tasks

语音生物标志物 | 6.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5201 字 阅读 →
论文解读

ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4331 字 阅读 →
论文解读

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

语音分离 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4749 字 阅读 →
论文解读

MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models

基准测试 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4488 字 阅读 →
论文解读

Modeling Both Intra- And Inter-Utterance Variability for Conversational Emotion Recognition

语音情感识别 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4447 字 阅读 →
论文解读

MSANET: Multi-Scale Semantic Aggregation Network for Brain-Assisted Speech Enhancement in Multi-Speaker Conditions

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4444 字 阅读 →
论文解读

MSCT: Differential Cross-Modal Attention for Deepfake Detection

音频深度伪造检测 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3968 字 阅读 →
论文解读

MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5146 字 阅读 →
论文解读

Multimodal Fusion-Based IPCLIP Network for Mixed Reality Surgical Assistance

多模态模型 | 6.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3263 字 阅读 →
论文解读

Multimodal LLMs as Expert Speech Annotators: Acoustic Macro-Descriptors for Parkinson's Detection

语音生物标志物 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4079 字 阅读 →
论文解读

Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4714 字 阅读 →
论文解读

Multimodal Transformer with Multiperspective Training for Predicting Self-Expression Skills from Video Interview

多模态模型 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4288 字 阅读 →
论文解读

MusiCRS: Benchmarking Audio-Centric Conversational Recommendation

音乐推荐 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4258 字 阅读 →
论文解读

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5908 字 阅读 →
论文解读

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

多模态模型 | 8.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6861 字 阅读 →
论文解读

Non-Line-of-Sight Vehicle Detection via Audio-Visual Fusion

音频分类 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4086 字 阅读 →
论文解读

OMNI-AVSR: Towards Unified Multimodal Speech Recognition With Large Language Models

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4649 字 阅读 →
论文解读

Perceptual Quality Assessment for Stylized Talking Heads

模型评估 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4192 字 阅读 →
论文解读

PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos

歌唱语音合成 | 4.5/10

 · 更新于 2026-09-25 · 约 3 分钟 · 1471 字 阅读 →