论文解读

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

音频字幕生成 | 8.4/10

 · 更新于 2026-09-25 · 约 24 分钟 · 11629 字 阅读 →
论文解读

VIOLET: High-Fidelity Violin Synthesis with Techniques and Dynamics

音乐生成 | 7.6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7116 字 阅读 →
论文解读

Assessing AI-generated music detection in real-world broadcast monitoring

音频伪造检测 | 6.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6180 字 阅读 →
论文解读

Frame-Level Pansori Mode Classification with Complementary Audio Representations

音乐理解 | 7.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6284 字 阅读 →
论文解读

MMAG: A Multi-Control Mixed Audio Generation Benchmark

音频生成 | 6.8/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7097 字 阅读 →
论文解读

Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset

音乐转录 | 7.7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5987 字 阅读 →
论文解读

C^3PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models

音视频问答 | 5.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7600 字 阅读 →
论文解读

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model

音乐生成 | 6.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6469 字 阅读 →
论文解读

Crowdsourced Multilingual Speech Intelligibility Testing

语音质量评估 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5390 字 阅读 →
论文解读

Identity-Faithful Audio-Visual Target Speaker Extraction with QIANGDA and VOXBLINK2-AVSE

音视频语音分离 | 6.8/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10746 字 阅读 →
论文解读

CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

音视频生成 | 7.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5209 字 阅读 →
论文解读

SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

语音识别 | 8.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5919 字 阅读 →
论文解读

Learning to Predict Performance-induced Emotion Differences in Classical Piano Music

数据集 | 5.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5773 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-08-03

共分析 16 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 46 分钟 · 22753 字 阅读 →
论文解读

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

音视频理解 | 8.3/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5406 字 阅读 →
论文解读

CrowdioSet and PaRIRset: Two Datasets Towards Live Music Source Separation

音乐源分离 | 7.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5803 字 阅读 →
论文解读

A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

说话人日志 | 8.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6772 字 阅读 →
论文解读

CARE: A Multimodal Corpus for Studying Speech and Non-Verbal Communication Across Multiple Medical Conditions

音视频理解 | 7.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8510 字 阅读 →
论文解读

Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance

语音识别 | 7.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5808 字 阅读 →
论文解读

Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis

语音合成 | 6.9/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10244 字 阅读 →