论文解读

HCGAN: Harmonic-Coupled Generative Adversarial Network for Speech Super-Resolution in Low-Bandwidth Scenarios

语音增强 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4284 字 阅读 →
论文解读

HVAC-EAR: Eavesdropping Human Speech Using HVAC Systems

音频安全 | 8.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5225 字 阅读 →
论文解读

HyFlowSE: Hybrid End-To-End Flow-Matching Speech Enhancement via Generative-Discriminative Learning

语音增强 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4024 字 阅读 →
论文解读

Improving Contextual Asr Via Multi-Grained Fusion With Large Language Models

语音识别 | 8.5/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3976 字 阅读 →
论文解读

Joint Autoregressive Modeling of Multi-Talker Overlapped Speech Recognition and Translation

语音识别 语音翻译 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4119 字 阅读 →
论文解读

Joint Deep Secondary Path Estimation and Adaptive Control for Active Noise Cancellation

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4814 字 阅读 →
论文解读

Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-Task Multi-Scale Network

音乐理解 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4795 字 阅读 →
论文解读

K-Function: Joint Pronunciation Transcription and Feedback for Evaluating Kids Language Function

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4671 字 阅读 →
论文解读

Language-Infused Retrieval-Augmented CTC with Adaptive Soft-Hard Gating for Robust Code-Switching ASR

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3170 字 阅读 →
论文解读

Lattice-Guided Consistency Regularization of Dual-Mode Transducers for Automatic Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 8 分钟 · 3973 字 阅读 →
论文解读

Learning to Align with Unbalanced Optimal Transport in Linguistic Knowledge Transfer for ASR

语音识别 | 6.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4273 字 阅读 →
论文解读

Lightweight Implicit Neural Network for Binaural Audio Synthesis

空间音频 | 7.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6286 字 阅读 →
论文解读

Lingometer: On-Device Personal Speech Word Counting System

语音活动检测 | 8.0/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5825 字 阅读 →
论文解读

Low-Bandwidth High-Fidelity Speech Transmission with Generative Latent Joint Source-Channel Coding

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4889 字 阅读 →
论文解读

MELA-TTS: Joint Transformer-Diffusion Model with Representation Alignment for Speech Synthesis

语音合成 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5124 字 阅读 →
论文解读

Mixture of Experts for Recognizing Depression from Interview and Reading Tasks

语音生物标志物 | 6.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5201 字 阅读 →
论文解读

MSANET: Multi-Scale Semantic Aggregation Network for Brain-Assisted Speech Enhancement in Multi-Speaker Conditions

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4444 字 阅读 →
论文解读

Musicdetr: A Position-Aware Spectral Note Detection Model for Singing Transcription

歌唱语音转录 | 8.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4470 字 阅读 →
论文解读

Peeking Into the Future for Contextual Biasing

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5086 字 阅读 →
论文解读

Polynomial Mixing for Efficient Self-Supervised Speech Encoders

语音识别 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4653 字 阅读 →