论文解读

FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations

语音合成 | 6.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 7010 字 阅读 →
论文解读

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio

语音合成 | 10/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6556 字 阅读 →
论文解读

LongCat-Video-Avatar 1.5 Technical Report

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7471 字 阅读 →
论文解读

PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech

语音合成 | 6.5/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8339 字 阅读 →
论文解读

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis

语音合成 | 9.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7303 字 阅读 →
论文解读

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation

语音合成 | 7.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6239 字 阅读 →
论文解读

Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control

语音合成 | 6.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5960 字 阅读 →
论文解读

WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5927 字 阅读 →
论文解读

Continual Speaker Identity Unlearning with Minimal Interference

语音合成 | 8.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6104 字 阅读 →
论文解读

CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS

语音合成 | 8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6277 字 阅读 →
论文解读

FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations

语音合成 | 8.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6292 字 阅读 →
论文解读

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio

音频水印 | 6.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6460 字 阅读 →
论文解读

Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control

语音合成 | 5.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5339 字 阅读 →
论文解读

WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models

语音合成 | 9.4/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5426 字 阅读 →
论文解读

Natural Yet Challenging to Detect: Robust In-the-Wild TTS through EMA and Dual-Scoring Prompt Selection -- Submission for WildSpoof 2026 TTS Track

语音合成 | 5.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5445 字 阅读 →
论文解读

StepAudio 2.5 Technical Report

统一音频模型 | 8.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6187 字 阅读 →
论文解读

Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition

语音质量评估 | 8.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7861 字 阅读 →
论文解读

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech

语音合成 | 9.5/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9678 字 阅读 →
论文解读

Bridging the Gap: Converting Read Text to Conversational Dialogue

语音转换 | 3.1/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5452 字 阅读 →
论文解读

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis

语音合成 | 6.8/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7239 字 阅读 →