论文解读

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

音频场景理解 | 8.0/10

 · 更新于 2026-09-17 · 约 10 分钟 · 4842 字 阅读 →
论文解读

Spectrographic Portamento Gradient Analysis: A Quantitative Method for Historical Cello Recordings with Application to Beethoven's Piano and Cello Sonatas, 1930--2012

音乐信息检索 | 7.5/10

 · 更新于 2026-09-17 · 约 7 分钟 · 3281 字 阅读 →
论文解读

Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations

音乐信息检索 | 8.0/10

 · 更新于 2026-09-17 · 约 9 分钟 · 4294 字 阅读 →
论文解读

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

语音质量评估 | 7.5/10

 · 更新于 2026-09-17 · 约 12 分钟 · 5555 字 阅读 →
论文解读

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

音频生成 | 8.5/10

 · 更新于 2026-09-17 · 约 12 分钟 · 5972 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-04-27

共分析 13 篇语音/AI 论文

 · 更新于 2026-09-17 · 约 41 分钟 · 20102 字 阅读 →
论文解读

MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

语音合成 | 7.5/10

 · 更新于 2026-09-17 · 约 8 分钟 · 3859 字 阅读 →
论文解读

MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation

机器人技能学习 | 7.5/10

 · 更新于 2026-09-17 · 约 8 分钟 · 3742 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-04-25

共分析 2 篇语音/AI 论文

 · 更新于 2026-09-17 · 约 6 分钟 · 2986 字 阅读 →
论文解读

"This Wasn't Made for Me": Recentering User Experience and Emotional Impact in the Evaluation of ASR Bias

语音识别 | 7.0/10

 · 更新于 2026-09-17 · 约 4 分钟 · 1997 字 阅读 →
论文解读

ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis

语音合成 | 7.0/10

 · 更新于 2026-09-17 · 约 14 分钟 · 6680 字 阅读 →
论文解读

AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA

音频问答 | 6.5/10

 · 更新于 2026-09-17 · 约 5 分钟 · 2444 字 阅读 →
论文解读

Beyond Rules: Towards Basso Continuo Personal Style Identification

音乐理解 | 7.0/10

 · 更新于 2026-09-17 · 约 7 分钟 · 3212 字 阅读 →
论文解读

DiariZen Explained: A Tutorial for the Open Source State-of-the-Art Speaker Diarization Pipeline

说话人分离 | 6.5/10

 · 更新于 2026-09-17 · 约 10 分钟 · 4881 字 阅读 →
论文解读

Dilated CNNs for Periodic Signal Processing: A Low-Complexity Approach

语音增强 | 6.5/10

 · 更新于 2026-09-17 · 约 5 分钟 · 2311 字 阅读 →
论文解读

Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-17 · 约 10 分钟 · 4532 字 阅读 →
论文解读

Evaluation of Automatic Speech Recognition Using Generative Large Language Models

语音识别 | 7.5/10

 · 更新于 2026-09-17 · 约 7 分钟 · 3189 字 阅读 →
论文解读

Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge

语音对话系统 | 6.5/10

 · 更新于 2026-09-17 · 约 7 分钟 · 3206 字 阅读 →
论文解读

Hierarchical Policy Optimization for Simultaneous Translation of Unbounded Speech

语音翻译 | 7.5/10

 · 更新于 2026-09-17 · 约 8 分钟 · 3626 字 阅读 →
论文解读

Low-Rank Adaptation Redux for Large Models

大语言模型 | 5.5/10

 · 更新于 2026-09-17 · 约 4 分钟 · 1870 字 阅读 →