论文解读

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring

多模态模型 | 7.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5751 字 阅读 →
论文解读

SceneBind: Binding What and Where Across Vision, Audio and Language

音视频理解 | 6.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6403 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-17

共分析 15 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 52 分钟 · 25719 字 阅读 →
论文解读

ARIMA: Reconstruction-Grounded Predictive Representation Learning for Symbolic Music

自监督学习 | 7.7/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7552 字 阅读 →
论文解读

Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment

音乐生成 | 7.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6661 字 阅读 →
论文解读

Local Multimodal Music Alignment from Global Supervision

对比学习 | 7.6/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7732 字 阅读 →
论文解读

MRUF: Multi-granularity Routing with Uncertainty-Aware Fusion for Robust Multimodal Sentiment Analysis

多模态模型 | 5.9/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8734 字 阅读 →
论文解读

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition

音视频语音识别 | 6.9/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7585 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-13

共分析 14 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 51 分钟 · 25294 字 阅读 →
论文解读

COALA: Robust Contextualized Speech-augmented Language Modeling for ASR via Contrastive Regularizer and Biasing Score Estimation

语音识别 | 8.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6706 字 阅读 →
论文解读

COALA: Robust Contextualized Speech-augmented Language Modeling for ASR via Contrastive Regularizer and Biasing Score Estimation

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9228 字 阅读 →
论文解读

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

音乐理解 | 8.1/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12029 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-09

共分析 13 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 47 分钟 · 23491 字 阅读 →
论文解读

Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking

音乐检索 | 5.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6172 字 阅读 →
论文解读

Uncovering Latent Depression Severity for Binary Depression Detection via Advantage-weighting Ranking

音视频理解 | 7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6394 字 阅读 →
论文解读

\(C^3\)ASD: Multi-Level Consistency-Driven Representation Learning

音视频理解 | 7.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7018 字 阅读 →
论文解读

Doppelganger: Sound Effects and Their Synthetic Twins

音频检索 | 9.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6213 字 阅读 →
论文解读

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment

语音情感识别 | 5.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7228 字 阅读 →
论文解读

Bioacoustic Geolocation: Species Sounds as Geographic Signals

音频理解 | 5.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6365 字 阅读 →
论文解读

Efficient Multi-modal Dataset Distillation via Analytic Parameter Matching

对比学习 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6217 字 阅读 →