论文解读

Mental Damage: Caption Poisoning Attacks on Retrieval-Augmented Text-to-Music Generation

音乐生成 | 5.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5959 字 阅读 →
论文解读

MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors

语音合成 | 8.5/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8301 字 阅读 →
论文解读

ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood

语音识别 | 8/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5167 字 阅读 →
论文解读

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

语音合成 | 8.6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6179 字 阅读 →
论文解读

The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models

语音识别 | 7/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6823 字 阅读 →
论文解读

Building Community-Centred NLP Resources for Puno Quechua

语音识别 | 7.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5330 字 阅读 →
论文解读

A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

多模态模型 | 7.7/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6784 字 阅读 →
论文解读

Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy

自监督学习 | 8.1/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6023 字 阅读 →
论文解读

Continual Speaker Identity Unlearning with Minimal Interference

语音合成 | 8.3/10

 · 更新于 2026-09-06 · 约 4 分钟 · 1998 字 阅读 →
论文解读

LongCat-Video-Avatar 1.5 Technical Report

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7471 字 阅读 →
论文解读

MERIT: Learning Disentangled Music Representations for Audio Similarity

音频检索 | 9/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5351 字 阅读 →
论文解读

Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5492 字 阅读 →
论文解读

A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

语音情感识别 | 7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6117 字 阅读 →
论文解读

Multilingual Phonological Feature Recognition with Self-Supervised Speech Models

语音识别 | 7.7/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5589 字 阅读 →
论文解读

MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio

音频深度伪造检测 | 10/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7102 字 阅读 →
论文解读

In Silico Modeling of the RAMPHO Buffer: Dissociating Informational and Energetic Masking via Phonetic Entropy in Deep Neural Networks

认知科学 | 6.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5140 字 阅读 →
论文解读

EnvTriCascade: An Environment-Aware Tri-Stage Cascaded Framework for ESDD2 2026 Challenge

音频深度伪造检测 | 5.3/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9344 字 阅读 →
论文解读

SAME: A Semantically-Aligned Music Autoencoder

音频编码 | 8.5/10

 · 更新于 2026-09-06 · 约 26 分钟 · 13010 字 阅读 →
论文解读

Toward World Modeling of Physiological Signals with Chaos-Theoretic Balancing and Latent Dynamics

生理信号预测 | 6/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8109 字 阅读 →
论文解读

AudioMosaic: Contrastive Masked Audio Representation Learning

音频分类 | 7.3/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7392 字 阅读 →