论文解读

UniVocal: Unified Speech-Singing Code-Switching Synthesis

语音合成 | 8.9/10

 · 更新于 2026-09-25 · 约 5 分钟 · 2097 字 阅读 →
论文解读

AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing

音乐生成 | 8.6/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6772 字 阅读 →
论文解读

Escaping the Linearity Trap: Manifold Detours for Black-Box Adversarial Attacks on Singing Audio Deepfake Detection

自监督学习 | 9.7/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7568 字 阅读 →
论文解读

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

语音合成 | 9.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6263 字 阅读 →
论文解读

MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors

语音合成 | 8.5/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8301 字 阅读 →
论文解读

ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood

语音识别 | 8/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5167 字 阅读 →
论文解读

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

语音合成 | 8.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6179 字 阅读 →
论文解读

The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models

语音识别 | 7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6823 字 阅读 →
论文解读

Building Community-Centred NLP Resources for Puno Quechua

语音识别 | 7.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5330 字 阅读 →
论文解读

A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

多模态模型 | 7.7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6784 字 阅读 →
论文解读

Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy

自监督学习 | 8.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6023 字 阅读 →
论文解读

Continual Speaker Identity Unlearning with Minimal Interference

语音合成 | 8.3/10

 · 更新于 2026-09-25 · 约 4 分钟 · 1998 字 阅读 →
论文解读

LongCat-Video-Avatar 1.5 Technical Report

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7471 字 阅读 →
论文解读

MERIT: Learning Disentangled Music Representations for Audio Similarity

音频检索 | 9/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5351 字 阅读 →
论文解读

Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5492 字 阅读 →
论文解读

A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

语音情感识别 | 7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6117 字 阅读 →
论文解读

Multilingual Phonological Feature Recognition with Self-Supervised Speech Models

语音识别 | 7.7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5589 字 阅读 →
论文解读

MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio

音频深度伪造检测 | 10/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7102 字 阅读 →
论文解读

In Silico Modeling of the RAMPHO Buffer: Dissociating Informational and Energetic Masking via Phonetic Entropy in Deep Neural Networks

认知科学 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5140 字 阅读 →
论文解读

EnvTriCascade: An Environment-Aware Tri-Stage Cascaded Framework for ESDD2 2026 Challenge

音频深度伪造检测 | 5.3/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9344 字 阅读 →