论文解读

ViP-VL: Vietnamese Self-supervised Speech Pretraining Model with Vector-Quantization Learning

语音识别 | 9.7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5467 字 阅读 →
论文解读

A Comparative Study of Pre-trained Speech Encoders and Training Objectives for Large-Scale Indic Spoken Language Identification

自监督学习 | 8.9/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4667 字 阅读 →
论文解读

A Comparison of SSL-Based Feature Extractors and Back-End Classifiers for Spoofing Detection: A Multi-Corpus Training and Cross-Linguistic Analysis

自监督学习 | 5/10

 · 更新于 2026-09-06 · 约 22 分钟 · 10637 字 阅读 →
论文解读

BareWave: Waveform-Native Flow-Matching Text-to-Speech

语音合成 | 7.0/10

 · 更新于 2026-09-06 · 约 17 分钟 · 8292 字 阅读 →
论文解读

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis

语音合成 | 9/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6999 字 阅读 →
论文解读

Multi-View Speech Representation Learning for Parkinson's Disease Detection Using Context-guided Cross-modal Attention

自监督学习 | 7.9/10

 · 更新于 2026-09-06 · 约 12 分钟 · 6004 字 阅读 →
论文解读

Probing Token Spaces under Generator Shift in AI-Generated Music Detection

音频编码 | 9/10

 · 更新于 2026-09-06 · 约 12 分钟 · 6007 字 阅读 →
论文解读

Assessing True Generalisability of Audio-Visual Speech Recognisers

语音识别 | 9.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6841 字 阅读 →
论文解读

Geometric Second-Order Feature Correlation Learning for Self-Supervised Speech Emotion Recognition

语音情感识别 | 7.9/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5626 字 阅读 →
论文解读

HybridCodec: Fast Dual-Stream, Semantically Enhanced Neural Audio Codec

语音合成 | 5.7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6148 字 阅读 →
论文解读

Leveraging Soft Distributions of SSL-Derived Discrete Speech Tokens for Downstream Inference

语音识别 | 7.4/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6365 字 阅读 →
论文解读

Mitigating Proxy-to-Wild Domain Gap in Deepfake Speech

数据增强 | 8.4/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6757 字 阅读 →
论文解读

Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations

语音合成 | 8.4/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5643 字 阅读 →
论文解读

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5847 字 阅读 →
论文解读

TargetSEC: Plug-and-Play In-the-Wild Speech Emotion Conversion via Arousal-Conditioned Latent Style Diffusion

语音转换 | 6.8/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5649 字 阅读 →
论文解读

Age-Aware Adapter Tuning for Children's Speech Recognition

语音识别 | 8.4/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5531 字 阅读 →
论文解读

Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5073 字 阅读 →
论文解读

CoSTA: Cognitive-State-Conditioned TTS Data Augmentation Using ASR Transcripts for Alzheimer's Disease Detection

语音合成 | 6.5/10

 · 更新于 2026-09-06 · 约 4 分钟 · 1509 字 阅读 →
论文解读

Do speech foundation models perceive speaker similarity as humans do?

说话人识别 | 6.3/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6169 字 阅读 →
论文解读

Domain-Aware Mispronunciation Detection and Diagnosis Using Language-Specific Statistical Graphs

图神经网络 | 6.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5824 字 阅读 →