论文解读

MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5146 字 阅读 →
论文解读

Multimodal Fusion-Based IPCLIP Network for Mixed Reality Surgical Assistance

多模态模型 | 6.5/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3263 字 阅读 →
论文解读

Multimodal LLMs as Expert Speech Annotators: Acoustic Macro-Descriptors for Parkinson's Detection

语音生物标志物 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4079 字 阅读 →
论文解读

Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4714 字 阅读 →
论文解读

Multimodal Transformer with Multiperspective Training for Predicting Self-Expression Skills from Video Interview

多模态模型 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4288 字 阅读 →
论文解读

MusiCRS: Benchmarking Audio-Centric Conversational Recommendation

音乐推荐 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4258 字 阅读 →
论文解读

Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generation

音频生成 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5908 字 阅读 →
论文解读

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

多模态模型 | 8.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6861 字 阅读 →
论文解读

Non-Line-of-Sight Vehicle Detection via Audio-Visual Fusion

音频分类 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4086 字 阅读 →
论文解读

OMNI-AVSR: Towards Unified Multimodal Speech Recognition With Large Language Models

语音识别 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4649 字 阅读 →
论文解读

Perceptual Quality Assessment for Stylized Talking Heads

模型评估 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4192 字 阅读 →
论文解读

PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos

歌唱语音合成 | 4.5/10

 · 更新于 2026-09-06 · 约 3 分钟 · 1471 字 阅读 →
论文解读

Phrased: Phrase Dictionary Biasing for Speech Translation

语音翻译 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4074 字 阅读 →
论文解读

Prompt-Guided Mixture-of-Experts for Robust Multimodal Sentiment Analysis with Missing Modalities

语音情感识别 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4723 字 阅读 →
论文解读

PromptSep: Generative Audio Separation Via Multimodal Prompting

语音分离 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3891 字 阅读 →
论文解读

Prototype-Guided Cross-Modal Contrastive Learning for Continual Audio-Visual Sound Separation

语音分离 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4252 字 阅读 →
论文解读

Rationale-Guided Learning for Multimodal Emotion Recognition

语音情感识别 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4768 字 阅读 →
论文解读

RCAL: Reinforced Cross-Modal Alignment for Multimodal Sentiment Analysis with Sparse Visual Frames

多模态模型 | 8.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4562 字 阅读 →
论文解读

Reading Between the Waves: Robust Topic Segmentation Using Inter-Sentence Audio Features

音频分类 | 7.0/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5113 字 阅读 →
论文解读

Reasoning Driven Captions to Assist Noise Robust Speech Emotion Recognition

语音情感识别 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4690 字 阅读 →