论文解读

An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis

语音合成 | 8.2/10

 · 更新于 2026-09-08 · 约 10 分钟 · 4584 字 阅读 →
论文解读

AP-GRPO: Anchor-Gated Phonetic Alignment with Policy Optimization for Pathological Speech Reconstruction

语音识别 | 7.4/10

 · 更新于 2026-09-08 · 约 14 分钟 · 6856 字 阅读 →
论文解读

ArtBoost: Synthetic Articulatory Data Augmentation for Acoustic-to-Articulatory Inversion

语音识别 | 6.5/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5884 字 阅读 →
论文解读

ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition

语音识别 | 8.3/10

 · 更新于 2026-09-08 · 约 10 分钟 · 4766 字 阅读 →
论文解读

AUDEDIT: Inversion-Free Text-Guided Editing with Pretrained Audio Flow Models

生成模型 | 7.8/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5877 字 阅读 →
论文解读

Beyond Artifacts: Towards Generalizable Synthetic Song Detection via Music-Intrinsic Features

音乐信息检索 | 8.4/10

 · 更新于 2026-09-08 · 约 12 分钟 · 6011 字 阅读 →
论文解读

Beyond Classification: A Cough Regression Benchmark for Respiratory Acoustic Foundation Models

音频事件检测 | 6/10

 · 更新于 2026-09-08 · 约 13 分钟 · 6382 字 阅读 →
论文解读

Bridging the SEA Gap: An Initial Benchmark for Neural Audio Codec-Synthesized Speech Deepfakes in South-East Asian Languages

语音合成 | 8.2/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5381 字 阅读 →
论文解读

Bridging the Usability Gap: Lessons from Interpreting Studies for Machine Interpreting Design

语音翻译 | 7.1/10

 · 更新于 2026-09-08 · 约 8 分钟 · 3574 字 阅读 →
论文解读

Closed-Loop Triplet Synergistic Generation for Long-Form Video

Closed-Loop Triplet Synergistic Generation for Long-Form Video

 · 更新于 2026-09-08 · 约 11 分钟 · 5392 字 阅读 →
论文解读

Confidence Score Guided Incremental and Speaker Adaptive Pseudo-Labeling for Semi-Supervised Elderly Speech Recognition

语音识别 | 7.2/10

 · 更新于 2026-09-08 · 约 13 分钟 · 6027 字 阅读 →
论文解读

Connecting Speech to Words through Images

无监督学习 | 7.1/10

 · 更新于 2026-09-08 · 约 13 分钟 · 6213 字 阅读 →
论文解读

CraBERT: Efficient Phoneme Encoder Pre-Training via Cascade Fusion of Subword Representations for Text-to-Speech

语音合成 | 7.5/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5796 字 阅读 →
论文解读

Data-Driven Decoding of Russell's Circumplex Model of Affect

语音情感识别 | 7.2/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5482 字 阅读 →
论文解读

DDPO-VC: Speaker De-Identification via Diffusion Denoising Policy Optimization

语音转换 | 6.5/10

 · 更新于 2026-09-08 · 约 19 分钟 · 9064 字 阅读 →
论文解读

Decoding while Adapting: Zero-Shot Online Speaker Adaptation via Audio-Textual Prompts for Elderly Speech Recognition

语音识别 | 6.8/10

 · 更新于 2026-09-08 · 约 14 分钟 · 6816 字 阅读 →
论文解读

Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection

课程学习 | 7.2/10

 · 更新于 2026-09-08 · 约 27 分钟 · 13307 字 阅读 →
论文解读

DuraMark: Duration-Embedded Watermarking in LLM-based TTS

生成模型 | 8.7/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5243 字 阅读 →
论文解读

Dynamic Prosody Prediction in LLM-based TTS for Improving Speaker Similarity

语音合成 | 7.6/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5199 字 阅读 →
论文解读

EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning

音频问答 | 6.1/10

 · 更新于 2026-09-08 · 约 22 分钟 · 10753 字 阅读 →