论文解读

DASH: Dual-View Self-Distillation with Multi-Layer Hidden Representations for Robust Speech Recognition

语音识别 | 6.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5965 字 阅读 →
论文解读

NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization

声源定位 | 7.8/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5653 字 阅读 →
论文解读

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs

语音合成 | 7.4/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7228 字 阅读 →
论文解读

Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6503 字 阅读 →
论文解读

Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4484 字 阅读 →
论文解读

Learning task-specific subspaces via interventional post-training of speech foundation models

自监督学习 | 6.2/10

 · 更新于 2026-09-06 · 约 21 分钟 · 10293 字 阅读 →
论文解读

Next-Turn: Duration-Aware Streaming Endpoint Detection via Time-to-Next-Speech-Onset Prediction

语音合成 | 7.9/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5858 字 阅读 →
论文解读

Perceptual compensation for tonal context in self-supervised speech models

语音识别 | 7.7/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5350 字 阅读 →
论文解读

PhASE-Flow: Phonetic-Conditioned Acoustic Flow Matching in SSL Representation Domain for Speech Enhancement

语音增强 | 7.6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6160 字 阅读 →
论文解读

ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition

语音识别 | 8.3/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4766 字 阅读 →
论文解读

Beyond Classification: A Cough Regression Benchmark for Respiratory Acoustic Foundation Models

音频事件检测 | 6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6382 字 阅读 →
论文解读

CraBERT: Efficient Phoneme Encoder Pre-Training via Cascade Fusion of Subword Representations for Text-to-Speech

语音合成 | 7.5/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5796 字 阅读 →
论文解读

From Physics to Representation: Audio Learning with Synthetic Pre-training via Procedural Generation

自监督学习 | 8.2/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5339 字 阅读 →
论文解读

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction

语音合成 | 6.8/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5195 字 阅读 →
论文解读

NVMOS: Non-Verbal Vocalization Quality Assessment in Speech

自监督学习 | 6.2/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4781 字 阅读 →
论文解读

Rhythm of the Deep: A Computational-Linguistic Test of Duality of Patterning in Sperm Whale Codas

自监督学习 | 8.5/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6634 字 阅读 →
论文解读

Robust Spoofed Speech Detection via Temporal Pyramid Modeling

音频深度伪造检测 | 6.7/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6597 字 阅读 →
论文解读

Spectro-Temporal Interference Confounds Phase Encoding in Spatial Audio Foundation Models

自监督学习 | 7.4/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5251 字 阅读 →
论文解读

From Self-Supervised Speech Models to Mixture-of-Experts for Robust Anti-Spoofing

自监督学习 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5096 字 阅读 →
论文解读

Mask, Sample, Revise: A Revisable CTMC Inference Stack for Guided Discrete Flow Matching Text-to-Speech

语音合成 | 6.8/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8765 字 阅读 →