论文解读

Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec

语音增强 | 7.1/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5494 字 阅读 →
论文解读

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

音频事件检测 | 6.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5766 字 阅读 →
论文解读

RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation

语音增强 | 6.6/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6600 字 阅读 →
论文解读

The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6748 字 阅读 →
论文解读

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

音视频生成 | 7.8/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7019 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-08-13

共分析 18 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 56 分钟 · 27887 字 阅读 →
论文解读

A Dataset and Benchmark for Optical Music Recognition of String Quartet Scores

端到端 | 8.5/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7765 字 阅读 →
论文解读

ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS

语音合成 | 6.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6433 字 阅读 →
论文解读

Beyond Dry References: Learning Relative Audio Effects Representations via Contrastive Distance Learning

音乐理解 | 7.7/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5536 字 阅读 →
论文解读

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

语音质量评估 | 7.1/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8060 字 阅读 →
论文解读

BiTSE: Binaural Target Speaker Extraction in Noisy Multi-Talker Environments for AR Glass Arrays

语音分离 | 5.0/10

 · 更新于 2026-09-06 · 约 20 分钟 · 9824 字 阅读 →
论文解读

DINO-A: Adapting Self-Distillation Vision Transformers to General Audio Representation Learning

音频分类 | 6.2/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7322 字 阅读 →
论文解读

DIY e-HandPan: A new DIY Low-Cost Handpan Interface based on Arduino and ESP32 Microcontrollers

音频交互 | 7.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6288 字 阅读 →
论文解读

DuplexWorld: Can voice agents help you get through the day?

语音交互 | 8.5/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8820 字 阅读 →
论文解读

Edge Phoneme Recognition for Children's Speech through Age-Aware Training

语音识别 | 6.3/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6776 字 阅读 →
论文解读

Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence

音视频生成 | 6.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6768 字 阅读 →
论文解读

In Defense of Using Worst-case Privacy Disclosure as Privacy Evaluation Metric of Voice Anonymization

说话人验证 | 7.1/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5631 字 阅读 →
论文解读

MAJEPPA: Morphing and Assessing in a Unified Piano Performance Space

音乐理解 | 7.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6358 字 阅读 →
论文解读

MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model

音乐生成 | 6.6/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7239 字 阅读 →
论文解读

Measuring Cross-Cultural Style Diffusion Through Era Classification: US and Korean Popular Music

音乐理解 | 8.3/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7089 字 阅读 →