论文解读

Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5073 字 阅读 →
论文解读

CoSTA: Cognitive-State-Conditioned TTS Data Augmentation Using ASR Transcripts for Alzheimer's Disease Detection

语音合成 | 6.5/10

 · 更新于 2026-09-25 · 约 4 分钟 · 1509 字 阅读 →
论文解读

Do speech foundation models perceive speaker similarity as humans do?

说话人识别 | 6.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6169 字 阅读 →
论文解读

Domain-Aware Mispronunciation Detection and Diagnosis Using Language-Specific Statistical Graphs

图神经网络 | 6.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5824 字 阅读 →
论文解读

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation

语音合成 | 7.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6113 字 阅读 →
论文解读

M2S-AVSR: Modality-aware Multi-view Self-supervised Representation for Robust Audio-Visual Speech Recognition

语音识别 | 9/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6425 字 阅读 →
论文解读

ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6641 字 阅读 →
论文解读

USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding

音频编码 | 9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 6009 字 阅读 →
论文解读

Channel-Oriented Design for EEG-to-Music Reconstruction

音乐生成 | 7.7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4877 字 阅读 →
论文解读

DetectZoo: A Unified Toolkit for AI-Generated Content Detection Across Text, Audio, and Image Modalities

多模态模型 | 9.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6131 字 阅读 →
论文解读

SURF: Separation via Unsupervised Remixing Flow

无监督学习 | 6.4/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5166 字 阅读 →
论文解读

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement

语音合成 | 8.6/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6116 字 阅读 →
论文解读

MoDAl: Self-Supervised Neural Modality Discovery via Decorrelation for Speech Neuroprosthesis

自监督学习 | 6.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5302 字 阅读 →
论文解读

SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment

语音识别 | 7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6619 字 阅读 →
论文解读

SpeakerCard-1M: An Evidence-Grounded Speaker Card Corpus for In-the-Wild Speaker Verification

说话人验证 | 7.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6193 字 阅读 →
论文解读

Stable Hybrid Cross-Attention Fusion for Audio-Visual Event Recognition

自监督学习 | 6.7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5944 字 阅读 →
论文解读

A Lightweight Slot-Attention Framework for Multi-Instrument Multi-Pitch Estimation

音乐信息检索 | 6.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6340 字 阅读 →
论文解读

Context-aware child-directed speech detection from long-form recordings

自监督学习 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5403 字 阅读 →
论文解读

Echo: A Joint-Embedding Predictive Architecture for Speaker Diarization and Speech Recognition in a Shared Latent Space

语音识别 | 7/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7107 字 阅读 →
论文解读

Privacy-preserving Prosody Representation Learning

自监督学习 | 4.9/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5261 字 阅读 →