论文解读

SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models

语音识别 | 6.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6373 字 阅读 →
论文解读

Genre Bias or Aesthetic Perception? Identifying and Mitigating Shortcut Learning in Music Evaluation

音乐理解 | 6.0/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6179 字 阅读 →
论文解读

MetaPerch: Learning from metadata for bioacoustics foundation models

音频分类 | 9.0/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10694 字 阅读 →
论文解读

DOA Estimation from One-Bit Magnitude-Only Measurements via Sign-Consistency Optimization

声源定位 | 5.1/10

 · 更新于 2026-09-25 · 约 18 分钟 · 9009 字 阅读 →
论文解读

Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction

语音伪造检测 | 6.1/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7419 字 阅读 →
论文解读

MRUF: Multi-granularity Routing with Uncertainty-Aware Fusion for Robust Multimodal Sentiment Analysis

多模态模型 | 5.9/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8734 字 阅读 →
论文解读

MusicMark: A Robust Generative Watermarking Framework for Music Generation

音频水印 | 7.3/10

 · 更新于 2026-09-25 · 约 18 分钟 · 8863 字 阅读 →
论文解读

Tight-Frame Reconstruction for Acoustic Intensity Estimation Using Cardioid Microphone Pairs

声源定位 | 6.8/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7054 字 阅读 →
论文解读

COALA: Robust Contextualized Speech-augmented Language Modeling for ASR via Contrastive Regularizer and Biasing Score Estimation

语音识别 | 8.2/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6706 字 阅读 →
论文解读

Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models

语音属性识别 | 7.8/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3485 字 阅读 →
论文解读

An Intervention-Based Framework for Shortcut Diagnosis in Spoofing Countermeasures

语音伪造检测 | 6.1/10

 · 更新于 2026-09-25 · 约 22 分钟 · 10952 字 阅读 →
论文解读

\(C^3\)ASD: Multi-Level Consistency-Driven Representation Learning

音视频理解 | 7.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7018 字 阅读 →
论文解读

Physiological Noise Augmentation Improves Non-Invasive Brain-to-Speech

语音识别 | 6/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7173 字 阅读 →
论文解读

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization

音视频理解 | 4.8/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6505 字 阅读 →
论文解读

Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio

音频水印 | 7/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6164 字 阅读 →
论文解读

Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits

语音识别 | 9.3/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6415 字 阅读 →
论文解读

Multimodal Fusion via Self-Consistent Task-Gradient Fields

鲁棒性 | 5.5/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7110 字 阅读 →
论文解读

SONAR: Spectral‑Contrastive Audio Residuals for Generalizable Deepfake Detection

语音伪造检测 | 7.8/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7289 字 阅读 →
论文解读

Stable Spectral Copula Alignment for Robust Multimodal Learning

鲁棒性 | 5.2/10

 · 更新于 2026-09-25 · 约 15 分钟 · 7374 字 阅读 →
论文解读

Speaker head orientation estimation with a single microphone array using phase spectrogram features

声源定位 | 5.8/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5887 字 阅读 →