论文解读

Speech Enhancement Based on Drifting Models

语音增强 | 7.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5266 字 阅读 →
论文解读

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling

语音合成 | 7.5/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6055 字 阅读 →
论文解读

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

语音合成评估 | 7.0/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6277 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-04-28

共分析 24 篇语音/AI 论文

 · 更新于 2026-09-10 · 约 69 分钟 · 34252 字 阅读 →
论文解读

Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus

语音识别 | 7.0/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5316 字 阅读 →
论文解读

Audio Effect Estimation with DNN-Based Prediction and Search Algorithm

音乐理解 | 8.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4821 字 阅读 →
论文解读

Audio Video Verbal Analysis (AVVA) for Capturing Classroom Dialogues

音频问答 | 6.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4018 字 阅读 →
论文解读

Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis

发音错误检测 | 8.5/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6458 字 阅读 →
论文解读

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models

说话人识别 | 8.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4588 字 阅读 →
论文解读

Earable Platform with Integrated Simultaneous EEG Sensing and Auditory Stimulation

音频事件检测 | 5.5/10

 · 更新于 2026-09-10 · 约 6 分钟 · 2648 字 阅读 →
论文解读

Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge

语音对话系统 | 6.5/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3585 字 阅读 →
论文解读

Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models

语音识别 | 7.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4567 字 阅读 →
论文解读

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

音频场景理解 | 8.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4842 字 阅读 →
论文解读

Spectrographic Portamento Gradient Analysis: A Quantitative Method for Historical Cello Recordings with Application to Beethoven's Piano and Cello Sonatas, 1930--2012

音乐信息检索 | 7.5/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3281 字 阅读 →
论文解读

Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations

音乐信息检索 | 8.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4294 字 阅读 →
论文解读

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

语音质量评估 | 7.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5555 字 阅读 →
论文解读

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

音频生成 | 8.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5972 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-04-27

共分析 13 篇语音/AI 论文

 · 更新于 2026-09-10 · 约 41 分钟 · 20102 字 阅读 →
论文解读

MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control

语音合成 | 7.5/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3859 字 阅读 →
论文解读

MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation

机器人技能学习 | 7.5/10

 · 更新于 2026-09-10 · 约 8 分钟 · 3742 字 阅读 →