论文解读

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing

语音编辑 | 7.3/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6259 字 阅读 →
论文解读

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation

语音合成 | 7.0/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7896 字 阅读 →
论文解读

HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models

语音情感识别 | 7.2/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7528 字 阅读 →
论文解读

Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds

音频生成 | 7.0/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8583 字 阅读 →
论文解读

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

音频生成 | 7.1/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7166 字 阅读 →
论文解读

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

语音合成 | 6.0/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5892 字 阅读 →
论文解读

Phylogenetic signal in marine mammal and bird vocalizations captured by audio foundation models: the limited benefit of domain-specific pretraining

音频理解 | 7.7/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6153 字 阅读 →
论文解读

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis

语音合成 | 7.0/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7788 字 阅读 →
论文解读

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

语音识别 | 7.1/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8698 字 阅读 →
论文解读

FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration

音乐生成 | 7.9/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6828 字 阅读 →
论文解读

ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching

语音合成 | 7.3/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7717 字 阅读 →
论文解读

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance

音乐生成 | 6.8/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7974 字 阅读 →
论文解读

CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection

提示学习 | 8.8/10

 · 更新于 2026-09-24 · 约 21 分钟 · 10446 字 阅读 →
论文解读

MeloBottleneck: Self-Supervised Melody Skeleton Extraction with a Latent Subsequence Bottleneck

音乐理解 | 7.5/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8118 字 阅读 →
论文解读

Best-of-N TTS Evaluation is Confounded by ASR Family Alignment

语音合成 | 8.7/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9056 字 阅读 →
论文解读

RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain

音频理解 | 8.1/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6569 字 阅读 →
论文解读

TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion

语音转换 | 8/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5784 字 阅读 →
论文解读

MusicDET: Zero-Shot AI-Generated Music Detection

音频伪造检测 | 6.1/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7804 字 阅读 →
论文解读

Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration

音乐生成 | 6.5/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8684 字 阅读 →
论文解读

Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech

语音合成 | 7/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6988 字 阅读 →