论文解读

ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction

语音对话系统 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4678 字 阅读 →
论文解读

Scaling Speech Tokenizers with Diffusion Autoencoders

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3677 字 阅读 →
论文解读

SpeechJudge: Towards Human-Level Judgment for Speech Naturalness

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8929 字 阅读 →
论文解读

Toward Complex-Valued Neural Networks for Waveform Generation

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5338 字 阅读 →
论文解读

TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems

语音合成评估 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4384 字 阅读 →
论文解读

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

语音翻译 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4434 字 阅读 →
论文解读

VibeVoice: Expressive Podcast Generation with Next-Token Diffusion

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4511 字 阅读 →
论文解读

Continuous Audio Language Models

音频生成 音乐生成 | 9.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5470 字 阅读 →
论文解读

DrVoice: Parallel Speech-Text Voice Conversation Model via Dual-Resolution Speech Representations

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5139 字 阅读 →
论文解读

FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates

语音合成 | 8.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6115 字 阅读 →
论文解读

FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4459 字 阅读 →
论文解读

Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation

音频生成 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4234 字 阅读 →
论文解读

From Natural Alignment to Conditional Controllability in Multimodal Dialogue

语音合成 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4689 字 阅读 →
论文解读

Gogo: Group-wise granularity-ordered codec for stable and efficient speech generation

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4704 字 阅读 →
论文解读

Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech Synthesis

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6858 字 阅读 →
论文解读

JaiTTS: A Thai Voice Cloning Model

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6263 字 阅读 →
论文解读

Latent Speech-Text Transformer

语音大模型 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5696 字 阅读 →
论文解读

MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control

语音克隆 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4919 字 阅读 →
论文解读

ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction

语音对话系统 | 8.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6023 字 阅读 →
论文解读

SpeechJudge: Towards Human-Level Judgment for Speech Naturalness

模型评估 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5451 字 阅读 →