论文解读

Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations

音乐信息检索 | 8.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4287 字 阅读 →
论文解读

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

语音质量评估 | 7.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5548 字 阅读 →
论文解读

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

音频生成 | 8.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5965 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-05-01

共分析 21 篇语音/AI 论文

 · 更新于 2026-09-10 · 约 70 分钟 · 35064 字 阅读 →
论文解读

A New Location Estimator for Mixed LOS & NLOS scenarios

声源定位 | 7.5/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5838 字 阅读 →
论文解读

A Toolkit for Detecting Spurious Correlations in Speech Datasets

模型评估 | 7.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4619 字 阅读 →
论文解读

DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

语音匿名化 | 7.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4816 字 阅读 →
论文解读

Diffusion Reconstruction towards Generalizable Audio Deepfake Detection

音频深度伪造检测 | 7.5/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4407 字 阅读 →
论文解读

Dual-LoRA: Parameter-Efficient Adversarial Disentanglement for Cross-Lingual Speaker Verification

说话人验证 | 7.5/10

 · 更新于 2026-09-10 · 约 13 分钟 · 6131 字 阅读 →
论文解读

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses

语音情感识别 | 7.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5250 字 阅读 →
论文解读

Fitting Large Nonlinear Mixed Effects Models Using Variational Expectation Maximization

统计计算 | 6.5/10

 · 更新于 2026-09-10 · 约 6 分钟 · 2744 字 阅读 →
论文解读

Full band denoising of room impulse response in the wavelet domain with dictionary learning

音频信号处理 | 6.5/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4520 字 阅读 →
论文解读

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

音视频 | 8.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5319 字 阅读 →
论文解读

Hankel and Toeplitz Rank-1 Decomposition of Arbitrary Matrices with Applications to Signal Direction-of-Arrival Estimation

声源定位 | 7.5/10

 · 更新于 2026-09-10 · 约 7 分钟 · 3385 字 阅读 →
论文解读

Multimodal LLMs are not all you need for Pediatric Speech Language Pathology

语音分类 | 7.5/10

 · 更新于 2026-09-10 · 约 11 分钟 · 5063 字 阅读 →
论文解读

Multiple Additive Neural Networks for Structured and Unstructured Data

表格数据预测 | 7.0/10

 · 更新于 2026-09-10 · 约 10 分钟 · 4631 字 阅读 →
论文解读

One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech

语音克隆 | 8.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4425 字 阅读 →
论文解读

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

语音合成 | 8.0/10

 · 更新于 2026-09-10 · 约 12 分钟 · 5558 字 阅读 →
论文解读

Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

语音合成 | 7.5/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4165 字 阅读 →
论文解读

PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech

语音合成 | 9.0/10

 · 更新于 2026-09-10 · 约 9 分钟 · 4333 字 阅读 →