论文解读

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

语音合成 | 7.6/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7803 字 阅读 →
论文解读

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech

语音合成 | 7.9/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8042 字 阅读 →
论文解读

AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing

音频生成 | 5.8/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7545 字 阅读 →
论文解读

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

语音合成 | 8/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7969 字 阅读 →
论文解读

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability

语音合成 | 6.6/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6527 字 阅读 →
论文解读

FoeGlass: Simple In-Context Learning Is Enough for Red Teaming Audio Deepfake Detectors

语音伪造检测 | 6.8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6644 字 阅读 →
论文解读

Multimodal Latent Language Modeling with Next-Token Diffusion

语音合成 | 6.1/10

 · 更新于 2026-09-25 · 约 6 分钟 · 2865 字 阅读 →
论文解读

Optimality of FSQ Tokens for Continuous Diffusion for Categorical Data with Application to Text-to-Speech

语音合成 | 8/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5305 字 阅读 →
论文解读

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models

语音编码 | 8/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6645 字 阅读 →
论文解读

Scaling Transformers for End-to-End Discrete Audio Tokenization

音频编码 | 7.1/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5426 字 阅读 →
论文解读

Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech

语音合成 | 7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6988 字 阅读 →
论文解读

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

音视频生成 | 6.8/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7540 字 阅读 →
论文解读

Two-dimensional quantization for geometry-aware audio coding

语音编码 | 7.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5851 字 阅读 →
论文解读

Unlocking Speech–Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning

语音交互 | 6.7/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7653 字 阅读 →
论文解读

An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation

语音合成 | 6.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6302 字 阅读 →
论文解读

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings

语音合成 | 5.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5960 字 阅读 →
论文解读

Using embeddings to predict spoken word duration and pitch in Mandarin monosyllabic words

语音合成 | 4/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7644 字 阅读 →
论文解读

A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models

语音合成 | 6.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6328 字 阅读 →
论文解读

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation

语音合成 | 6.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6668 字 阅读 →
论文解读

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

语音合成 | 7.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6790 字 阅读 →