论文解读

Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances

音乐理解 | 7.7/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6828 字 阅读 →
论文解读

RIPPLE: Generating Multi-Channel Phase, Not Recovering It

音频理解 | 7.1/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8301 字 阅读 →
论文解读

SKY-Piano: A Multimodal Piano Performance Dataset

音频理解 | 8.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6136 字 阅读 →
论文解读

Teffic-Audio: Tell Fact from Fiction

语音伪造检测 | 6.8/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8390 字 阅读 →
论文解读

The MADRS Pipeline: Supporting Depression Assessment in Clinical Trials

医疗音频 | 5.8/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8722 字 阅读 →
论文解读

VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition

歌唱生成 | 7.5/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7140 字 阅读 →
论文解读

WeSep: A Modular and Cue-Composable Framework for Target Speaker Extraction

语音分离 | 7.4/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7245 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-31

共分析 16 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 53 分钟 · 26338 字 阅读 →
论文解读

A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

说话人日志 | 8.2/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6772 字 阅读 →
论文解读

A Study on Online Mask-based Beamforming Using Per-channel Masking for Spatially Distributed Microphones

语音增强 | 5.7/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4813 字 阅读 →
论文解读

Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection

语音伪造检测 | 6.6/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7231 字 阅读 →
论文解读

Detection of AI-generated stems within hybrid human-AI music

CNN | 6.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6505 字 阅读 →
论文解读

Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens

语音识别 | 6.3/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8991 字 阅读 →
论文解读

DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues

语音交互 | 6.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5720 字 阅读 →
论文解读

Explicit Note-Event Tokenization and Pitch-Validity Constrained Decoding for MIDI-to-Tablature Transcription

音乐转录 | 6.1/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5320 字 阅读 →
论文解读

Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement

音频分类 | 6.8/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7590 字 阅读 →
论文解读

Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

音频伪造检测 | 7.5/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7402 字 阅读 →
论文解读

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

音频字幕生成 | 6.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6369 字 阅读 →
论文解读

MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation

音乐生成 | 7.4/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4922 字 阅读 →
论文解读

Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis

音频交互 | 7.0/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6613 字 阅读 →