论文解读

SpeechOp: Inference-Time Task Composition for Generative Speech Processing

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5070 字 阅读 →
论文解读

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs

语音分词 | 9.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4686 字 阅读 →
论文解读

Toward Complex-Valued Neural Networks for Waveform Generation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4635 字 阅读 →
论文解读

TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems

模型评估 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4181 字 阅读 →
论文解读

VibeVoice: Expressive Podcast Generation with Next-Token Diffusion

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5927 字 阅读 →
论文解读

Accent Conversion: A Problem-Driven Survey of Sociolinguistic and Technical Constraints

语音转换 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3936 字 阅读 →
论文解读

Few-Shot Accent Synthesis for ASR with LLM-Guided Phoneme Editing

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5168 字 阅读 →
论文解读

JaiTTS: A Thai Voice Cloning Model

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2500 字 阅读 →
论文解读

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

语音质量评估 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5548 字 阅读 →
论文解读

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

音频生成 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5965 字 阅读 →
论文解读

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5250 字 阅读 →
论文解读

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5558 字 阅读 →
论文解读

Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4165 字 阅读 →
论文解读

PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech

语音合成 | 9.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4333 字 阅读 →
论文解读

SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4115 字 阅读 →
论文解读

ARCHI-TTS: A Flow-Matching-Based Text-to-Speech Model with Self-Supervised Semantic Aligner and Accelerated Inference

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6329 字 阅读 →
论文解读

Asynchrony-Aware Decoupled Multimodal Control for Cued Speech Video Generation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4846 字 阅读 →
论文解读

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5499 字 阅读 →
论文解读

Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4555 字 阅读 →
论文解读

BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5196 字 阅读 →