论文解读

OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains

数据增强 | 8.2/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6462 字 阅读 →
论文解读

Unsupervised Approaches for Global Prosodic Embedding Extraction

语音合成 | 7.8/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6086 字 阅读 →
论文解读

Decoding Insect Song: A Multitask Semisupervised Orthoptera Bioacoustic Classifier

音频分类 | 8.7/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6708 字 阅读 →
论文解读

Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations

音频分类 | 7.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5440 字 阅读 →
论文解读

From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial Animation

语音合成 | 7.9/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6037 字 阅读 →
论文解读

Leveraging Audio-LLMs to Filter Speech-to-Speech Training Data

语音翻译 | 8.4/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6061 字 阅读 →
论文解读

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment

语音合成 | 9.7/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6236 字 阅读 →
论文解读

Context-Aware Multimodal Claim Verification in Spoken Dialogues

多模态模型 | 7.1/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6922 字 阅读 →
论文解读

Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains

语音识别 | 9.1/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4665 字 阅读 →
论文解读

Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders

语音合成 | 7.7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5248 字 阅读 →
论文解读

Pretrained self-supervised speech models can recognize unseen consonants

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4929 字 阅读 →
论文解读

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

语音合成 | 7.9/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4632 字 阅读 →
论文解读

UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction

语音合成 | 8.1/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6102 字 阅读 →
论文解读

Automated Pronunciation Evaluation for Korean Toddler Speech using Speech Diarization and Self-Supervised Learning

说话人日志 | 6/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6726 字 阅读 →
论文解读

Multilingual Word-Level Forced Alignment with Self-Supervised Representations and Learned Dynamic Programming

自监督学习 | 6.3/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5292 字 阅读 →
论文解读

Overview of ESDD2: Environment-Aware Speech and Sound Deepfake Detection Challenge

数据增强 | 6.3/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8529 字 阅读 →
论文解读

Recovering the Zipfian Distribution in Unsupervised Term Discovery

自监督学习 | 8.7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5173 字 阅读 →
论文解读

Speaker Group Encoding in Self-supervised Speech Recognition Models

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5916 字 阅读 →
论文解读

SSL-GMMVC: Interpretable Voice Conversion via Locally Linear GMM Transforms in Self-Supervised Representation Space

语音转换 | 6.8/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5852 字 阅读 →
论文解读

Towards Robust Arabic Speech Emotion Recognition with Deep Learning

语音情感识别 | 6.4/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5530 字 阅读 →