论文解读

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead

语音合成 | 6.7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5591 字 阅读 →
论文解读

Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS

语音合成 | 7.2/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4932 字 阅读 →
论文解读

Word Lengthening as a Function of Utterance Position: A Multi-Corpus Study

语音合成 | 8.1/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4828 字 阅读 →
论文解读

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine

语音合成 | 7.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5736 字 阅读 →
论文解读

FlowEdit: Associative Memory for Lifelong Pronunciation Adaptation in Flow-Matching TTS

语音合成 | 10/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5749 字 阅读 →
论文解读

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

语音合成 | 7.7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5246 字 阅读 →
论文解读

Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech

语音识别 | 7.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5725 字 阅读 →
论文解读

Interpreting Content and Speaker Characteristics in Factorised Self-Supervised Subspaces

语音合成 | 5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4928 字 阅读 →
论文解读

Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations

语音合成 | 6.9/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7804 字 阅读 →
论文解读

Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning

语音识别 | 8.7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5408 字 阅读 →
论文解读

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

语音合成 | 5.7/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3893 字 阅读 →
论文解读

PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation

语音合成 | 7.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6816 字 阅读 →
论文解读

ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversion

语音合成 | 6.2/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7884 字 阅读 →
论文解读

Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation

语音合成 | 7.9/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6654 字 阅读 →
论文解读

Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning

语音合成 | 7.7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4674 字 阅读 →
论文解读

FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech

语音合成 | 7.6/10

 · 更新于 2026-09-25 · 约 23 分钟 · 11197 字 阅读 →
论文解读

GRIDEX: Grid-Grounded Forensic Explanations for Deepfake Spectrogram Analysis

语音合成 | 8.6/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7676 字 阅读 →
论文解读

MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data

语音合成 | 7.9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5712 字 阅读 →
论文解读

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

语音合成 | 7.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5547 字 阅读 →
论文解读

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs

语音合成 | 7.4/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7228 字 阅读 →