论文解读

Attacking UTMOS: Probing the Robustness of a Speech Quality Assessment Model

语音质量评估 | 8.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5401 字 阅读 →
论文解读

Two kinds of robustness are not the same: disentangling fault tolerance and low-SNR robustness in multi-domain event detection on real data

音频事件检测 | 8.9/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4694 字 阅读 →
论文解读

VeRe-Flow: Guiding Flow Matching toward Clean Speech via Velocity Contrastive Regularization and Representation Alignment for Noise-Robust Bandwidth Expansion

语音增强 | 7.7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4813 字 阅读 →
论文解读

Supervised Post-training of Speech Foundation Models for Robust Adaptation in Speech Deepfake Detection

语音伪造检测 | 7.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5991 字 阅读 →
论文解读

An Analysis of Untrained Deep Reservoir Networks for Audio Surveillance

音频事件检测 | 8.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5526 字 阅读 →
论文解读

CORTIS: Text-Only Adaptation of Spoken Language Models for Task-Oriented Voice Agents

语音识别 | 7.7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4555 字 阅读 →
论文解读

DASH: Dual-View Self-Distillation with Multi-Layer Hidden Representations for Robust Speech Recognition

语音识别 | 6.6/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5965 字 阅读 →
论文解读

NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization

声源定位 | 7.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5653 字 阅读 →
论文解读

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs

多模态模型 | 8.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5845 字 阅读 →
论文解读

Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech

语音合成 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5784 字 阅读 →
论文解读

VoxWatermark: A Large-Scale Benchmark for Audio Watermark Detection under Perturbations

鲁棒性 | 9.4/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4873 字 阅读 →
论文解读

MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7026 字 阅读 →
论文解读

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions

鲁棒性 | 7.1/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7369 字 阅读 →
论文解读

Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and Algorithm

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 21 分钟 · 10471 字 阅读 →
论文解读

MeanVC 2: Robust Low-Latency Streaming Zero-Shot Voice Conversion

语音转换 | 6.9/10

 · 更新于 2026-09-25 · 约 23 分钟 · 11480 字 阅读 →
论文解读

Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5073 字 阅读 →
论文解读

Beyond WER: A Paired Acoustic Stress Test for Ambient Clinical Scribes

语音识别 | 7.1/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5672 字 阅读 →
论文解读

SB-RF: Schrödinger Bridge Rectified Flow for One-Step Robust Speech Enhancement

语音增强 | 7.6/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5309 字 阅读 →
论文解读

Drift-Augmented Scoring: Text-Derived Noise Robustness for Zero-Shot Audio-Language Classification

音频分类 | 10/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4712 字 阅读 →
论文解读

Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy

语音识别 | 8.6/10

 · 更新于 2026-09-25 · 约 3 分钟 · 1214 字 阅读 →