论文解读

Beyond Acoustic Sparsity and Linguistic Bias: A Prompt-Free Paradigm for Mispronunciation Detection and Diagnosis

发音错误检测 | 8.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6451 字 阅读 →
论文解读

Few-Shot Accent Synthesis for ASR with LLM-Guided Phoneme Editing

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5168 字 阅读 →
论文解读

Selective Augmentation: Improving Universal Automatic Phonetic Transcription via G2P Bootstrapping

语音识别 | 6.0/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3160 字 阅读 →
论文解读

Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations

音乐信息检索 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4287 字 阅读 →
论文解读

Diffusion Reconstruction towards Generalizable Audio Deepfake Detection

音频深度伪造检测 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4407 字 阅读 →
论文解读

Multimodal LLMs are not all you need for Pediatric Speech Language Pathology

语音分类 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5063 字 阅读 →
论文解读

Text-Utilization for Encoder-dominated Speech Recognition Models

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 6 分钟 · 2856 字 阅读 →
论文解读

A Consistent Learning Depression Detection Framework Integrating Multi-View Attention

语音生物标志物 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4619 字 阅读 →
论文解读

A Framework for Controlled Multi-Speaker Audio Synthesis for Robustness Evaluation of Speaker Diarisation Systems

说话人日志 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4437 字 阅读 →
论文解读

A Metric Learning Approach to Heart Murmur Detection from Phonocardiogram Recordings

音频分类 | 7.7/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3727 字 阅读 →
论文解读

A Unsupervised Domain Adaptation Framework For Semi-Supervised Melody Extraction Using Confidence Matrix Replace and Nearest Neighbour Supervision

音乐信息检索 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4391 字 阅读 →
论文解读

Addressing Gradient Misalignment in Data-Augmented Training for Robust Speech Deepfake Detection

语音伪造检测 | 7.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3953 字 阅读 →
论文解读

Advancing Semi-Supervised Child Speech Recognition with Omni-Temporal Classification under Label Noise

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4912 字 阅读 →
论文解读

ATOM: Adaptive Token-Level Optimal Transport Mixup for Speech Translation

语音翻译 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4810 字 阅读 →
论文解读

Attentive Masked Self-Distillation for Respiratory Sound Classification

音频分类 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4934 字 阅读 →
论文解读

Automatic Music Sample Identification with Multi-Track Contrastive Learning

音频检索 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4382 字 阅读 →
论文解读

Auxiliary Multi-Label Training For Improving the Robustness of Audio Deepfake Detection on AI-Processed Data

音频深度伪造检测 | 6.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3885 字 阅读 →
论文解读

Cardiobridge-DM: Bridging Cross-Cohort Heart Sound Synthesis via Rhythm-Aware Semi-Supervised Diffusion

音频生成 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4823 字 阅读 →
论文解读

Content-Preserving Speech Representation Learning Via Adaptive Segment-Level Alignment

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5415 字 阅读 →
论文解读

Contrastive Timbre Representations for Musical Instrument And Synthesizer Retrieval

音频检索 | 7.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3791 字 阅读 →