论文解读

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation

语音合成 | 7.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7931 字 阅读 →
论文解读

LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6067 字 阅读 →
论文解读

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6593 字 阅读 →
论文解读

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

语音合成 | 7.3/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2255 字 阅读 →
论文解读

DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information

语音识别 | 8.9/10

 · 更新于 2026-09-25 · 约 35 分钟 · 17529 字 阅读 →
论文解读

FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars

语音合成 | 7.8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6765 字 阅读 →
论文解读

SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation

语音合成 | 6.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6319 字 阅读 →
论文解读

A Survey of Automated Presentation Coaching: Systems, Methods, and Open Challenges

语音识别 | 5.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6247 字 阅读 →
论文解读

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

语音合成 | 8.2/10

 · 更新于 2026-09-25 · 约 27 分钟 · 13336 字 阅读 →
论文解读

HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models

语音合成 | 6.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5997 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-29

共分析 16 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 46 分钟 · 22697 字 阅读 →
论文解读

Closing the Quality Gap in Low-Resource Text-to-Speech: LoRA Fine-Tuning of VoxCPM2 for Khmer and Korean

语音合成 | 6/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4968 字 阅读 →
论文解读

VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation

语音合成 | 7.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5290 字 阅读 →
论文解读

Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS

语音合成 | 7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5086 字 阅读 →
论文解读

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations

语音合成 | 5.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5965 字 阅读 →
论文解读

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs?

语音识别 | 7.1/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5097 字 阅读 →
论文解读

Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4894 字 阅读 →
论文解读

Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant

语音识别 | 6.2/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12144 字 阅读 →
论文解读

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis

语音合成 | 7.3/10

 · 更新于 2026-09-25 · 约 4 分钟 · 1666 字 阅读 →
论文解读

STEB: A Speech-to-Speech Translation Expressiveness Benchmark for Evaluating Beyond Translation Fidelity

语音翻译 | 7.8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6956 字 阅读 →