论文解读

SHB-AE: Spherical harmonic beamforming based Ambisonics encoding and upscaling method for smartphone microphone array

音频编码 | 6.7/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5593 字 阅读 →
论文解读

Beyond the Mouth: Upper-Face Affective Cues in Audiovisual Sentence Recognition under Acoustic Uncertainty

语音识别 | 5.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4920 字 阅读 →
论文解读

Improving acoustic drone detection generalization through pretraining and data augmentation

音频事件检测 | 7.7/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5246 字 阅读 →
论文解读

TARQ: Tail-Aware Reconstruction Quantization for Rare-Word Robust Automatic Speech Recognition

语音识别 | 10/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6908 字 阅读 →
论文解读

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio

音频水印 | 6.2/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6460 字 阅读 →
论文解读

Convex Low-resource Accent-Robust Language Detection in Speech Recognition

语音识别 | 7.5/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6301 字 阅读 →
论文解读

Natural Yet Challenging to Detect: Robust In-the-Wild TTS through EMA and Dual-Scoring Prompt Selection -- Submission for WildSpoof 2026 TTS Track

语音合成 | 5.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5445 字 阅读 →
论文解读

DASM: Domain-Aware Sharpness Minimization for Multi-Domain Voice Stream Steganalysis

音频隐写分析 | 7.4/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6558 字 阅读 →
论文解读

Executable Boundary Contracts for Sound Event Traces

音频事件检测 | 8.5/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9595 字 阅读 →
论文解读

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

语音识别 | 9.3/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7818 字 阅读 →
论文解读

Cross-Talk Speech Reduction, by Separation, for Separation

语音分离 | 8.3/10

 · 更新于 2026-09-25 · 约 20 分钟 · 9536 字 阅读 →
论文解读

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection

音频深度伪造检测 | 7.2/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8115 字 阅读 →
论文解读

Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation

语音识别 | 6.8/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10575 字 阅读 →
论文解读

When Vision Speaks for Sound

音视频 | 7.7/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9514 字 阅读 →
论文解读

A Fast Robust Adaptive filter using Improved Data-Reuse Method

声学回声消除 | 6.2/10

 · 更新于 2026-09-25 · 约 17 分钟 · 8222 字 阅读 →
论文解读

Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities

音频问答 | 6.5/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8983 字 阅读 →
论文解读

Fractional-Order Subband p-Norm Adaptive Filter via Transformation Nearest Kronecker Product Decomposition for Active Noise Control

自适应滤波 | 5/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8577 字 阅读 →
论文解读

Robust Soft-Constrained Spatially Selective Active Noise Control for Hearables Under Secondary Path Variations

音频增强 | 5.7/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7281 字 阅读 →
论文解读

Adaptive Diagonal Loading using Krylov Subspaces for Robust Beamforming

声源定位 | 7.5/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8591 字 阅读 →
论文解读

Latent Secret Spin: Keyed Orthogonal Rotations for Blind Speech Watermarking in Anisotropic Latent Spaces

音频水印 | 5.5/10

 · 更新于 2026-09-25 · 约 16 分钟 · 7659 字 阅读 →