论文解读

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis

语音合成 | 6.4/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7332 字 阅读 →
论文解读

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

语音合成 | 7.0/10

 · 更新于 2026-09-24 · 约 21 分钟 · 10289 字 阅读 →
论文解读

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

音视频理解 | 6.6/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6883 字 阅读 →
论文解读

PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue

语音交互 | 8.0/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6734 字 阅读 →
论文解读

SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation

音频生成 | 7.6/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9341 字 阅读 →
论文解读

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation

语音合成 | 7.0/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7896 字 阅读 →
论文解读

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks

语音伪造检测 | 8.4/10

 · 更新于 2026-09-24 · 约 24 分钟 · 11604 字 阅读 →
论文解读

EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot

音视频交互 | 7.1/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7757 字 阅读 →
论文解读

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces

语音合成 | 5.2/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6366 字 阅读 →
论文解读

GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech model

语音合成 | 8.0/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8648 字 阅读 →
论文解读

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation

歌唱生成 | 6.1/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6210 字 阅读 →
论文解读

Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling

语音合成 | 6.1/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6963 字 阅读 →
论文解读

Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation

语音合成 | 7.8/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6646 字 阅读 →
论文解读

Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study

语音合成 | 6.8/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6254 字 阅读 →
论文解读

Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech

语音合成 | 7.2/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5978 字 阅读 →
论文解读

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

语音合成 | 6.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5563 字 阅读 →
论文解读

Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry

音频伪造检测 | 7.2/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6916 字 阅读 →
论文解读

Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil

语音伪造检测 | 7.7/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7303 字 阅读 →
论文解读

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

语音合成 | 5.8/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7230 字 阅读 →
论文解读

VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition

歌唱生成 | 7.5/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7140 字 阅读 →