论文解读

Query-Based Asymmetric Modeling with Decoupled Input–Output Rates for Speech Restoration

语音增强 | 7.1/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6737 字 阅读 →
论文解读

Robust Signal Enhancement via Fractional Detail Views and Knowledge Guided Multi-view Fusion

语音增强 | 5.7/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8257 字 阅读 →
论文解读

LMPAN: A Lightweight Multi-Path Alignment Network for Joint Full-Duplex Acoustic Echo Cancellation and Noise Suppression

语音增强 | 6.2/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8743 字 阅读 →
论文解读

RT-Tango: Real-Time Distributed Binaural Speech Enhancement for Low-Power Hearing Aid Devices

语音增强 | 5.5/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6468 字 阅读 →
论文解读

Spatial Speech Perception Systems: A Survey of Sound Source Localization, Directional Enhancement, and Speech Recognition

声源定位 | 4.1/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5837 字 阅读 →
论文解读

BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations

语音识别 | 6.9/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4866 字 阅读 →
论文解读

Improving multichannel speech enhancement through accurate room-acoustic simulations

语音增强 | 6.8/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5624 字 阅读 →
论文解读

OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7900 字 阅读 →
论文解读

VeRe-Flow: Guiding Flow Matching toward Clean Speech via Velocity Contrastive Regularization and Representation Alignment for Noise-Robust Bandwidth Expansion

语音增强 | 7.7/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4813 字 阅读 →
论文解读

Advancing Speaker-Based Vocal Effort Classification with WavLM and Data Augmentation in Naturalistic Non-Calibrated Speech Recordings

语音增强 | 6.8/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4639 字 阅读 →
论文解读

A Large-Scale Database and Predictive Model of Listener-Rated Ease of Speech Understanding in Commercial Hearing Aids

语音质量评估 | 8.1/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5406 字 阅读 →
论文解读

Error-Aware TF-IDF Retrieval-Augmented Generation for ASR Error Correction

语音识别 | 6.1/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4514 字 阅读 →
论文解读

Joint Residual Reweighting for Classifier Free Guidance in Flow-Matching Zero-Shot TTS

语音合成 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4894 字 阅读 →
论文解读

One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications

语音增强 | 7.2/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5687 字 阅读 →
论文解读

SE-AGCNet: An End-to-End Framework for Joint Speech Enhancement and Loudness Control in Meeting Scenarios

语音增强 | 7.4/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5590 字 阅读 →
论文解读

A Variational-Flow Analysis of StoRM under Noise-Power Mismatch

语音增强 | 4.4/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4515 字 阅读 →
论文解读

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

语音增强 | 8.1/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4879 字 阅读 →
论文解读

Beyond U-Net: A Latent-Representation-Aligned Skip-Free Backbone for Flow-Matching Speech Enhancement

语音增强 | 6.6/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6019 字 阅读 →
论文解读

Joint Learning of Covariance Estimation and White Noise Gain for Robust MVDR Beamforming

语音增强 | 5.8/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4629 字 阅读 →
论文解读

Neuromorphic Speech Enhancement with Dual-Branch Spiking Neural Networks

语音增强 | 7.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5146 字 阅读 →