论文解读

Robust Accent Identification via Voice Conversion and Non-Timbral Embeddings

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3282 字 阅读 →
论文解读

RRPO: Robust Reward Policy Optimization for LLM-Based Emotional TTS

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4067 字 阅读 →
论文解读

SA-SSL-MOS: Self-Supervised Learning MOS Prediction with Spectral Augmentation for Generalized Multi-Rate Speech Assessment

语音质量评估 | 7.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5780 字 阅读 →
论文解读

Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models

语音情感识别 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4289 字 阅读 →
论文解读

SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper

语音识别 | 8.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5790 字 阅读 →
论文解读

Source Separation For A Cappella Music

语音分离 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4161 字 阅读 →
论文解读

Style Attack Disguise: When Fonts Become a Camouflage for Adversarial Intent

对抗样本 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4288 字 阅读 →
论文解读

SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4862 字 阅读 →
论文解读

Synthesized Data Selection via Score Distribution Matching for Te Reo Māori Automatic Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4219 字 阅读 →
论文解读

Synthetic Data Domain Adaptation for ASR via LLM-Based Text and Phonetic Respelling Augmentation

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4965 字 阅读 →
论文解读

Three Seconds is Sufficient: A Multi-Pronged Framework for Model-Based Speaker Adaptation in ASR Under Data-Scarce Conditions

语音识别 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4318 字 阅读 →
论文解读

Timbre-Aware Audio Difference Captioning for Anomalous Machine Sounds without Paired Training Data via Synthetic Perturbations

音频分类 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4605 字 阅读 →
论文解读

Tldiffgan: A Latent Diffusion-Gan Framework with Temporal Information Fusion for Anomalous Sound Detection

音频事件检测 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4622 字 阅读 →
论文解读

Towards Blind Data Cleaning: A Case Study in Music Source Separation

音乐信息检索 | 7.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4929 字 阅读 →
论文解读

Towards Distance-Aware Synthetic Audio Mixtures for Universal Sound Separation

语音分离 | 6.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3929 字 阅读 →
论文解读

Towards Effective Negation Modeling in Joint Audio-Text Models for Music

音乐理解 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4421 字 阅读 →
论文解读

Training-Free Inference-Time Scaling for Audio Source Separation

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3556 字 阅读 →
论文解读

UNMIXX: Untangling Highly Correlated Singing Voices Mixtures

语音分离 | 8.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4383 字 阅读 →
论文解读

Vioptt: Violin Technique-Aware Transcription from Synthetic Data Augmentation

音乐信息检索 | 6.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4985 字 阅读 →
论文解读

WAV2LEV: Predicting Levenshtein Edit Operation Sequences For Fine-Grained Estimation of Automatic Speech Recognition Error

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4224 字 阅读 →