论文解读

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation

语音合成 | 7.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6113 字 阅读 →
论文解读

M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition

语音识别 | 9/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6425 字 阅读 →
论文解读

ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity

语音情感识别 | 7.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6641 字 阅读 →
论文解读

USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding

音频编码 | 9/10

 · 更新于 2026-09-06 · 约 12 分钟 · 6009 字 阅读 →
论文解读

Channel-Oriented Design for EEG-to-Music Reconstruction

音乐生成 | 7.7/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4877 字 阅读 →
论文解读

DetectZoo: A Unified Toolkit for AI-Generated Content Detection Across Text, Audio, and Image Modalities

多模态模型 | 9.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6131 字 阅读 →
论文解读

SURF: Separation via Unsupervised Remixing Flow

无监督学习 | 6.4/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5166 字 阅读 →
论文解读

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement

语音合成 | 8.6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6116 字 阅读 →
论文解读

MoDAl: Self-Supervised Neural Modality Discovery via Decorrelation for Speech Neuroprosthesis

自监督学习 | 6.6/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5302 字 阅读 →
论文解读

SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment

语音识别 | 7/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6619 字 阅读 →
论文解读

SpeakerCard-1M: An Evidence-Grounded Speaker Card Corpus for In-the-Wild Speaker Verification

说话人验证 | 7.4/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6193 字 阅读 →
论文解读

Stable Hybrid Cross-Attention Fusion for Audio-Visual Event Recognition

自监督学习 | 6.7/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5944 字 阅读 →
论文解读

A Lightweight Slot-Attention Framework for Multi-Instrument Multi-Pitch Estimation

音乐信息检索 | 6.7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6340 字 阅读 →
论文解读

Context-aware child-directed speech detection from long-form recordings

自监督学习 | 8.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5403 字 阅读 →
论文解读

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space

语音识别 | 7/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7107 字 阅读 →
论文解读

Privacy-preserving Prosody Representation Learning

自监督学习 | 4.9/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5261 字 阅读 →
论文解读

UniVocal: Unified Speech-Singing Code-Switching Synthesis

语音合成 | 8.9/10

 · 更新于 2026-09-06 · 约 5 分钟 · 2097 字 阅读 →
论文解读

AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing

音乐生成 | 8.6/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6772 字 阅读 →
论文解读

Escaping the Linearity Trap: Manifold Detours for Black-Box Adversarial Attacks on Singing Audio Deepfake Detection

自监督学习 | 9.7/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7568 字 阅读 →
论文解读

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

语音合成 | 9.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6263 字 阅读 →