论文解读

DSRMS-TransUnet: A Decentralized Non-Shifted Transunet for Shallow Water Acoustic Source Range Estimation

声源定位 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4613 字 阅读 →
论文解读

Dual-Strategy-Enhanced Conbimamba for Neural Speaker Diarization

说话人分离 | 8.0/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4642 字 阅读 →
论文解读

E2E-AEC: Implementing An End-To-End Neural Network Learning Approach for Acoustic Echo Cancellation

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4163 字 阅读 →
论文解读

EEND-SAA: Enrollment-Less Main Speaker Voice Activity Detection Using Self-Attention Attractors

语音活动检测 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4559 字 阅读 →
论文解读

Exploring SSL Discrete Tokens for Multilingual Automatic Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5356 字 阅读 →
论文解读

HAVT-IVD: Heterogeneity-Aware Cross-Modal Network for Audio-Visual Surveillance: Idling Vehicles Detection with Multichannel Audio and Multiscale Visual Cues

音频事件检测 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4420 字 阅读 →
论文解读

HCGAN: Harmonic-Coupled Generative Adversarial Network for Speech Super-Resolution in Low-Bandwidth Scenarios

语音增强 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4284 字 阅读 →
论文解读

HVAC-EAR: Eavesdropping Human Speech Using HVAC Systems

音频安全 | 8.5/10

 · 更新于 2026-09-06 · 约 11 分钟 · 5225 字 阅读 →
论文解读

HyFlowSE: Hybrid End-To-End Flow-Matching Speech Enhancement via Generative-Discriminative Learning

语音增强 | 8.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4024 字 阅读 →
论文解读

Improving Contextual Asr Via Multi-Grained Fusion With Large Language Models

语音识别 | 8.5/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3976 字 阅读 →
论文解读

Joint Autoregressive Modeling of Multi-Talker Overlapped Speech Recognition and Translation

语音识别 语音翻译 | 7.0/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4119 字 阅读 →
论文解读

Joint Deep Secondary Path Estimation and Adaptive Control for Active Noise Cancellation

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4814 字 阅读 →
论文解读

Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-Task Multi-Scale Network

音乐理解 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4795 字 阅读 →
论文解读

K-Function: Joint Pronunciation Transcription and Feedback for Evaluating Kids Language Function

语音识别 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4671 字 阅读 →
论文解读

Language-Infused Retrieval-Augmented CTC with Adaptive Soft-Hard Gating for Robust Code-Switching ASR

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 7 分钟 · 3170 字 阅读 →
论文解读

Lattice-Guided Consistency Regularization of Dual-Mode Transducers for Automatic Speech Recognition

语音识别 | 8.0/10

 · 更新于 2026-09-06 · 约 8 分钟 · 3973 字 阅读 →
论文解读

Learning to Align with Unbalanced Optimal Transport in Linguistic Knowledge Transfer for ASR

语音识别 | 6.5/10

 · 更新于 2026-09-06 · 约 9 分钟 · 4273 字 阅读 →
论文解读

Lightweight Implicit Neural Network for Binaural Audio Synthesis

空间音频 | 7.0/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6286 字 阅读 →
论文解读

Lingometer: On-Device Personal Speech Word Counting System

语音活动检测 | 8.0/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5825 字 阅读 →
论文解读

Low-Bandwidth High-Fidelity Speech Transmission with Generative Latent Joint Source-Channel Coding

语音增强 | 7.5/10

 · 更新于 2026-09-06 · 约 10 分钟 · 4889 字 阅读 →