论文解读

An Analysis of the Effectiveness of Synthetic Speech Data for ASR Fine-tuning in Selected Indic Languages

语音识别 | 5.7/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12144 字 阅读 →
论文解读

DeSRPA: Decoupled Speech Role-Playing Agent via Inference-Time Intervention

语音合成 | 7.3/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5315 字 阅读 →
论文解读

JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence

语音合成 | 7.7/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4464 字 阅读 →
论文解读

Next-Turn: Duration-Aware Streaming Endpoint Detection via Time-to-Next-Speech-Onset Prediction

语音合成 | 7.9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5858 字 阅读 →
论文解读

One-Step Token-to-Waveform Generation with MeanFlow in Latent Space

语音合成 | 9.3/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5638 字 阅读 →
论文解读

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI

语音识别 | 7.6/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6566 字 阅读 →
论文解读

An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis

语音合成 | 8.2/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4584 字 阅读 →
论文解读

Bridging the SEA Gap: An Initial Benchmark for Neural Audio Codec-Synthesized Speech Deepfakes in South-East Asian Languages

语音合成 | 8.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5381 字 阅读 →
论文解读

Dynamic Prosody Prediction in LLM-based TTS for Improving Speaker Similarity

语音合成 | 7.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5199 字 阅读 →
论文解读

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction

语音合成 | 6.8/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5195 字 阅读 →
论文解读

Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5784 字 阅读 →
论文解读

BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM

语音合成 | 9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5563 字 阅读 →
论文解读

FoleyGenEx: Unified Video-to-Audio Generation with Multi-Modal Control, Temporal Alignment, and Semantic Precision

语音合成 | 7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6238 字 阅读 →
论文解读

Mask, Sample, Revise: A Revisable CTMC Inference Stack for Guided Discrete Flow Matching Text-to-Speech

语音合成 | 6.8/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8765 字 阅读 →
论文解读

Unsupervised Approaches for Global Prosodic Embedding Extraction

语音合成 | 7.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6086 字 阅读 →
论文解读

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech

语音合成 | 9.3/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5378 字 阅读 →
论文解读

From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial Animation

语音合成 | 7.9/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6037 字 阅读 →
论文解读

Low-Latency Real-Time Audio Game Commentary System via LLM-Based Parallel Text Generation

语音合成 | 7.9/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4803 字 阅读 →
论文解读

NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation

语音翻译 | 7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4630 字 阅读 →
论文解读

PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue

语音合成 | 8.1/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10641 字 阅读 →