论文解读

AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4850 字 阅读 →
论文解读

FlexiVoice: Enabling Flexible Style Control in Zero-Shot TTS with Natural Language Instructions

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5075 字 阅读 →
论文解读

Learning multimodal dictionary decompositions with group-sparse autoencoders

跨模态检索 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3971 字 阅读 →
论文解读

SpeechOp: Inference-Time Task Composition for Generative Speech Processing

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4967 字 阅读 →
论文解读

SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization

音频检索 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4249 字 阅读 →
论文解读

Tell me Habibi, is it Real or Fake?

音视频深度伪造检测 | 8.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3642 字 阅读 →
论文解读

VibeVoice: Expressive Podcast Generation with Next-Token Diffusion

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4511 字 阅读 →
论文解读

AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5530 字 阅读 →
论文解读

DiffSDA: Unsupervised Diffusion Sequential Disentanglement Across Modalities

无监督学习 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4931 字 阅读 →
论文解读

Gogo: Group-wise granularity-ordered codec for stable and efficient speech generation

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4704 字 阅读 →
论文解读

Hierarchical Semantic-Acoustic Modeling via Semi-Discrete Residual Representations for Expressive End-to-End Speech Synthesis

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6858 字 阅读 →
论文解读

Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis

发音错误检测 | 8.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6451 字 阅读 →
论文解读

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation

语音情感识别 模型评估 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5334 字 阅读 →
论文解读

Affect-Jigsaw: Integrating Core and Peripheral Emotions for Harmonious Fine-Grained Multimodal Emotion Recognition

语音情感识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4583 字 阅读 →
论文解读

ARCHI-TTS: A Flow-Matching-Based Text-to-Speech Model with Self-Supervised Semantic Aligner and Accelerated Inference

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6329 字 阅读 →
论文解读

BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5196 字 阅读 →
论文解读

Conditional Diffusion Models for Mental Health-Preserving Voice Conversion

语音转换 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4927 字 阅读 →
论文解读

Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis

语音克隆 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4837 字 阅读 →
论文解读

DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis

语音合成 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5316 字 阅读 →
论文解读

Detecting and Attributing Synthetic Spanish Speech: The HISPASpoof Dataset

语音伪造检测 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4286 字 阅读 →