论文解读

Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems

语音识别 | 7.5/10

 · 更新于 2026-10-01 · 约 11 分钟 · 5492 字 阅读 →
论文解读

Rubato: Transcribing Piano Music with Timestamps

音乐转录 | 10/10

 · 更新于 2026-10-01 · 约 18 分钟 · 8956 字 阅读 →
论文解读

Score-Agnostic Structure Analysis in Large-Scale Performance Datasets

音乐信息检索 | 6.5/10

 · 更新于 2026-10-01 · 约 11 分钟 · 5291 字 阅读 →
论文解读

Subspace Track-before-Detect for Passive Multi-Target Tracking with Unknown Emitted Signals

信号处理基础 | 6.4/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6144 字 阅读 →
论文解读

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation

语音合成 | 7.7/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6239 字 阅读 →
论文解读

Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization

语音识别 | 6.0/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6387 字 阅读 →
论文解读

Time Segmented Beamforming via Dynamic Programming: Theory and Implementation

自适应滤波 | 8/10

 · 更新于 2026-10-01 · 约 14 分钟 · 6669 字 阅读 →
论文解读

Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control

语音合成 | 6.3/10

 · 更新于 2026-10-01 · 约 12 分钟 · 5960 字 阅读 →
论文解读

Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction

语音编码 | 8.1/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6291 字 阅读 →
论文解读

WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models

语音合成 | 8.5/10

 · 更新于 2026-10-01 · 约 12 分钟 · 5927 字 阅读 →
论文解读

Why Can't They Remember? Uncovering Representation and Retrieval Bottlenecks in Multi-Turn Acoustic Memory

语音识别 | 7/10

 · 更新于 2026-10-01 · 约 4 分钟 · 1800 字 阅读 →
论文解读

Zero-Shot Parkinson's Disease Detection from Speech: Comparing Large Audio and Language Models

大语言模型 | 5.2/10

 · 更新于 2026-10-01 · 约 14 分钟 · 6816 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-27

共分析 39 篇语音/AI 论文

 · 更新于 2026-10-01 · 约 121 分钟 · 60546 字 阅读 →
论文解读

A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

语音情感识别 | 7/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6117 字 阅读 →
论文解读

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

音频生成 | 7/10

 · 更新于 2026-10-01 · 约 10 分钟 · 4951 字 阅读 →
论文解读

Continual Speaker Identity Unlearning with Minimal Interference

语音合成 | 8.6/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6104 字 阅读 →
论文解读

CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS

语音合成 | 8/10

 · 更新于 2026-10-01 · 约 13 分钟 · 6277 字 阅读 →
论文解读

cSTMM: A Unified Complex Spherical Student's \(t\) Mixture Model for Directional Statistics in Mask-Based Blind Speech Separation

语音分离 | 7.9/10

 · 更新于 2026-10-01 · 约 12 分钟 · 5715 字 阅读 →
论文解读

Decoding Stimulus Reconstruction-Based Auditory Attention Robustly in Unbalanced EEG Datasets

交叉验证 | 8.9/10

 · 更新于 2026-10-01 · 约 16 分钟 · 7665 字 阅读 →
论文解读

Exploration of Perceptual Speech Features for Clinical Decision-Support in Mental Health Care

语音情感识别 | 8.9/10

 · 更新于 2026-10-01 · 约 15 分钟 · 7043 字 阅读 →