论文解读

wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2

自监督学习 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5028 字 阅读 →
论文解读

A Survey of Automated Presentation Coaching: Systems, Methods, and Open Challenges

语音识别 | 5.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6247 字 阅读 →
论文解读

Advancing Speaker-Based Vocal Effort Classification with WavLM and Data Augmentation in Naturalistic Non-Calibrated Speech Recordings

语音增强 | 6.8/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4639 字 阅读 →
论文解读

Do Speech Emphasis Models Generalize across Languages and Emotions?

语音识别 | 7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4935 字 阅读 →
论文解读

HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models

语音合成 | 6.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5997 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-29

共分析 16 篇语音/AI 论文

 · 更新于 2026-09-25 · 约 46 分钟 · 22697 字 阅读 →
论文解读

wav2tok 2.0: Scalable Audio Tokenization Maintaining Explicit Pairwise Token Alignment for Efficient Audio Retrieval

语音检索 | 6.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5876 字 阅读 →
论文解读

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users

端到端 | 7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6231 字 阅读 →
论文解读

Frequency-Aware Self-Supervised Music Representation Learning

音乐信息检索 | 6.8/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9505 字 阅读 →
论文解读

Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4894 字 阅读 →
论文解读

MJEPA: A Simple and Scalable Joint-Embedding Predictive Architecture for Audio-Visual Learning

自监督学习 | 7.4/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5126 字 阅读 →
论文解读

Phoneme-Level Mispronunciation Screening in Polish-Speaking Children with an Explainable Assistant

语音识别 | 6.2/10

 · 更新于 2026-09-25 · 约 25 分钟 · 12144 字 阅读 →
论文解读

Supervised Post-training of Speech Foundation Models for Robust Adaptation in Speech Deepfake Detection

语音伪造检测 | 7.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5991 字 阅读 →
论文解读

What Does a Pathological Speech Assessment Model Know about Acoustic Features? A Case Study on Oral and Oropharyngeal Cancer Patients

语音可懂度评估 | 6.4/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5414 字 阅读 →
论文解读

A Fusion-Aware Two-Stage Framework for Mispronunciation Detection and Diagnosis in Low-Resource Modern Standard Arabic

语音识别 | 7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5029 字 阅读 →
论文解读

Audio--Image Alignment as a Continued-Pretraining Stage Improves Low-Resource ASR

语音识别 | 6.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6034 字 阅读 →
论文解读

Autoencoder based optimized SSL representations: Complexity Minimization and improved Dysarthric ASR

语音识别 | 5.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4861 字 阅读 →
论文解读

Acoustic Landmark Detector based on Conformer and HuBERT

语音识别 | 5.5/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10238 字 阅读 →
论文解读

Backdoor Attacks on Speech Emotion Recognition via TTS-Generated Poisoning

语音情感识别 | 7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5489 字 阅读 →
论文解读

Bridging Self-Supervised Learning and Speech Enhancement: A Wav2Vec2-Conditioned Framework

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5269 字 阅读 →