论文解读

Mind Your [m]S, Cross Your [t]S: a Large-Scale Phonetic Analysis of Speech Reproduction in Modern Speech Generators

语音伪造检测 | 7.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4597 字 阅读 →
论文解读

MirrorTalk: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion Control

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3679 字 阅读 →
论文解读

Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation

说话人日志 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4023 字 阅读 →
论文解读

NCF-TTS: Enhancing Flow Matching Based Text-To-Speech with Neighborhood Consistency Flow

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4168 字 阅读 →
论文解读

Neuromamba: Adaptive Frequency Filtering with a Pyramid Mamba for sEEG-driven Speech Synthesis

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5962 字 阅读 →
论文解读

No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4301 字 阅读 →
论文解读

Optimizing Speech Language Models for Acoustic Consistency

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4747 字 阅读 →
论文解读

OV-INSTRUCTTTS: Towards Open-Vocabulary Instruct Text-to-Speech

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5743 字 阅读 →
论文解读

PFluxTTS: Hybrid Flow-Matching TTS with Robust Cross-Lingual Voice Cloning and Inference-Time Model Fusion

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6124 字 阅读 →
论文解读

Phonological Tokenizer: Prosody-Aware Phonetic Token Via Multi-Objective Fine-Tuning with Differentiable K-Means

语音表示学习 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5452 字 阅读 →
论文解读

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4562 字 阅读 →
论文解读

Principled Coarse-Grained Acceptance For Speculative Decoding In Speech

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3911 字 阅读 →
论文解读

Prosody-Guided Harmonic Attention for Phase-Coherent Neural Vocoding in the Complex Spectrum

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4866 字 阅读 →
论文解读

PRSA: Preventing Malicious Speaker Recognition and Speech Synthesis Simultaneously with Adversarial Examples

语音匿名化 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4445 字 阅读 →
论文解读

PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech

基准测试 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5099 字 阅读 →
论文解读

PSTalker: Realistic 3D Talking Head Synthesis via a Semantic-Aware Audio-Driven Point-Based Shape

说话人合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4433 字 阅读 →
论文解读

QFOCUS: Controllable Synthesis for Automated Speech Stress Editing to Deliver Human-Like Emphatic Intent

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 4 分钟 · 1766 字 阅读 →
论文解读

Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4134 字 阅读 →
论文解读

Real-Time Streaming MEL Vocoding with Generative Flow Matching

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4120 字 阅读 →
论文解读

Residual Tokens Enhance Masked Autoencoders for Speech Modeling

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5021 字 阅读 →