论文解读

Building an ASR Solution for Training and Assessing Children's Reading

语音识别 | 8.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5557 字 阅读 →
论文解读

VeRe-Flow: Guiding Flow Matching toward Clean Speech via Velocity Contrastive Regularization and Representation Alignment for Noise-Robust Bandwidth Expansion

语音增强 | 7.7/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4813 字 阅读 →
论文解读

CORTIS: Text-Only Adaptation of Spoken Language Models for Task-Oriented Voice Agents

语音识别 | 7.7/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4555 字 阅读 →
论文解读

Gradient-Based Learning of Parametric Engine Sound Representations for Real-Time Resynthesis and Tuning on Embedded Systems

参数高效微调 | 7.8/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4491 字 阅读 →
论文解读

STAR-VAE: Structured Topology-Aware Regularization for Audio Reconstruction and Generation

音频生成 | 8.8/10

 · 更新于 2026-09-24 · 约 24 分钟 · 11842 字 阅读 →
论文解读

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs

语音合成 | 7.4/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7228 字 阅读 →
论文解读

Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 9 分钟 · 4484 字 阅读 →
论文解读

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment

语音合成 | 9.7/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6236 字 阅读 →
论文解读

Frozen Multimodal Embeddings for Personality and Cognitive Ability Assessment in Asynchronous Video Interviews

语音情感识别 | 6.7/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6098 字 阅读 →
论文解读

Lung-SRAD: Spectral-Aware Regularized Audio DASS with Dual-Axis Patch-Mix Contrastive Learning for Respiratory Sound Classification

对比学习 | 6.8/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5479 字 阅读 →
论文解读

Quality Adaptive Angular Margin Learning for Respiratory Sound Classification

音频质量评估 | 9.5/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8931 字 阅读 →
论文解读

Profy: Interpretable Visualization of Expertise-Dependent Motor Skills Toward Supporting Piano Practice

音乐信息检索 | 6.9/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5822 字 阅读 →
论文解读

Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading

语音合成 | 7.5/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7716 字 阅读 →
论文解读

Discovering Functionally Selective Brain Regions with a Deep Topographic Multimodal Model

多模态模型 | 8.2/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5464 字 阅读 →
论文解读

MeanVC 2: Robust Low-Latency Streaming Zero-Shot Voice Conversion

语音转换 | 6.9/10

 · 更新于 2026-09-24 · 约 23 分钟 · 11480 字 阅读 →
论文解读

DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation

语音合成 | 7.5/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6978 字 阅读 →
论文解读

Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5492 字 阅读 →