论文解读

Recovering the Zipfian Distribution in Unsupervised Term Discovery

自监督学习 | 8.7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5173 字 阅读 →
论文解读

Speaker Group Encoding in Self-supervised Speech Recognition Models

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5916 字 阅读 →
论文解读

SSL-GMMVC: Interpretable Voice Conversion via Locally Linear GMM Transforms in Self-Supervised Representation Space

语音转换 | 6.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5852 字 阅读 →
论文解读

Towards Robust Arabic Speech Emotion Recognition with Deep Learning

语音情感识别 | 6.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5530 字 阅读 →
论文解读

ViP-VL: Vietnamese Self-supervised Speech Pretraining Model with Vector-Quantization Learning

语音识别 | 9.7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5467 字 阅读 →
论文解读

A Comparative Study of Pre-trained Speech Encoders and Training Objectives for Large-Scale Indic Spoken Language Identification

自监督学习 | 8.9/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4667 字 阅读 →
论文解读

A Comparison of SSL-Based Feature Extractors and Back-End Classifiers for Spoofing Detection: A Multi-Corpus Training and Cross-Linguistic Analysis

自监督学习 | 5/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10637 字 阅读 →
论文解读

BareWave: Waveform-Native Flow-Matching Text-to-Speech

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8292 字 阅读 →
论文解读

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis

语音合成 | 9/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6999 字 阅读 →
论文解读

Multi-View Speech Representation Learning for Parkinson's Disease Detection Using Context-guided Cross-modal Attention

自监督学习 | 7.9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 6004 字 阅读 →
论文解读

Probing Token Spaces under Generator Shift in AI-Generated Music Detection

音频编码 | 9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 6007 字 阅读 →
论文解读

Assessing True Generalisability of Audio-Visual Speech Recognisers

语音识别 | 9.5/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6841 字 阅读 →
论文解读

Geometric Second-Order Feature Correlation Learning for Self-Supervised Speech Emotion Recognition

语音情感识别 | 7.9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5626 字 阅读 →
论文解读

HybridCodec: Fast Dual-Stream, Semantically Enhanced Neural Audio Codec

语音合成 | 5.7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6148 字 阅读 →
论文解读

Leveraging Soft Distributions of SSL-Derived Discrete Speech Tokens for Downstream Inference

语音识别 | 7.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6365 字 阅读 →
论文解读

Mitigating Proxy-to-Wild Domain Gap in Deepfake Speech

数据增强 | 8.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6757 字 阅读 →
论文解读

Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations

语音合成 | 8.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5643 字 阅读 →
论文解读

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5847 字 阅读 →
论文解读

TargetSEC: Plug-and-Play In-the-Wild Speech Emotion Conversion via Arousal-Conditioned Latent Style Diffusion

语音转换 | 6.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5649 字 阅读 →
论文解读

Age-Aware Adapter Tuning for Children's Speech Recognition

语音识别 | 8.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5531 字 阅读 →