论文解读

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

音视频生成 | 6.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6768 字 阅读 →
论文解读

MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model

音乐生成 | 6.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7239 字 阅读 →
论文解读

Rationale-Guided Learning for Multimodal Emotion Recognition

语音情感识别 | 7.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7248 字 阅读 →
论文解读

Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition

语音识别 | 6.1/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6535 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-08-12

共分析 27 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 76 分钟 · 37591 字 阅读 →
论文解读

Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs

模型剪枝 | 8.4/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7856 字 阅读 →
论文解读

From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios

音视频语音识别 | 6.6/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7993 字 阅读 →
论文解读

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

多模态模型 | 5.9/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7905 字 阅读 →
论文解读

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

音视频问答 | 6.7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5213 字 阅读 →
论文解读

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

音视频理解 | 6.6/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6883 字 阅读 →
论文解读

Multilingual Emotion Neurons in Large Audio-Language Models

语音情感识别 | 6.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7199 字 阅读 →
论文解读

omni-macos: On-Device Omni-Modal Search on Apple Silicon

音频检索 | 7.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6772 字 阅读 →
论文解读

RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction

音频生成 | 6.9/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9065 字 阅读 →
论文解读

REFRAMED: Towards Realistic Audio Description Generation for Movies

音频字幕生成 | 8.0/10

 · 更新于 2026-09-25 · 约 18 分钟 · 9013 字 阅读 →
论文解读

SAMOT: State-Aware Step Modulation and Optimal Transport Matching for Audio-Visual Instance Segmentation

音视频理解 | 7.7/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12060 字 阅读 →
论文解读

Steering dense music retrieval with open-vocabulary concept discovery

音乐检索 | 7.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6289 字 阅读 →
论文解读

Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

语音属性识别 | 6.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6013 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-08-11

共分析 40 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 114 分钟 · 56810 字 阅读 →
论文解读

Objects as Audio-Visual Modal Sound Fields

音频生成 | 7.6/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6923 字 阅读 →
论文解读

Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation

音视频交互 | 5.7/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8051 字 阅读 →