论文解读

Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in Wav2vec 2.0

语音质量评估 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4296 字 阅读 →
论文解读

TinyMU: A Compact Audio-Language Model for Music Understanding

音乐理解 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5011 字 阅读 →
论文解读

TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for Ü-Tsang, Amdo and Kham Speech Dataset Generation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4601 字 阅读 →
论文解读

Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER

语音识别 | 9.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4930 字 阅读 →
论文解读

Training Dynamics-Aware Multi-Factor Curriculum Learning for Target Speaker Extraction

语音分离 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4427 字 阅读 →
论文解读

Training Flow Matching Models with Reliable Labels via Self-Purification

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4189 字 阅读 →
论文解读

Unsupervised Discovery and Analysis of the Vocal Repertoires and Patterns of Select Corvid Species

生物声学 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4918 字 阅读 →
论文解读

UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model

语音对话系统 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4909 字 阅读 →
论文解读

Visual Keys to Symphonies: Latent Diffusion for Multi-Scene Video-to-Music Generation

音乐生成 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5068 字 阅读 →
论文解读

ViTex: Visual Texture Control for Multi-Track Symbolic Music Generation via Discrete Diffusion Models

音乐生成 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3841 字 阅读 →
论文解读

WAV2LEV: Predicting Levenshtein Edit Operation Sequences For Fine-Grained Estimation of Automatic Speech Recognition Error

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4224 字 阅读 →
论文解读

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

音频场景理解 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4909 字 阅读 →
论文解读

RTCFake: Speech Deepfake Detection in Real-Time Communication

语音伪造检测 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4522 字 阅读 →
论文解读

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

语音合成评估 | 7.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6277 字 阅读 →
论文解读

Listening with Time: Precise Temporal Awareness for Long-Form Audio Understanding

音频场景理解 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4842 字 阅读 →
论文解读

Spectrographic Portamento Gradient Analysis: A Quantitative Method for Historical Cello Recordings with Application to Beethoven's Piano and Cello Sonatas, 1930--2012

音乐信息检索 | 7.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3281 字 阅读 →
论文解读

AUDITA: A New Dataset to Audit Humans vs. AI Skill at Audio QA

音频问答 | 6.5/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2444 字 阅读 →
论文解读

Beyond Rules: Towards Basso Continuo Personal Style Identification

音乐理解 | 7.0/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3212 字 阅读 →
论文解读

Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge

语音对话系统 | 6.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3206 字 阅读 →
论文解读

Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0

语音生物标志物 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4210 字 阅读 →