论文解读

Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction

语音伪造检测 | 6.1/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7419 字 阅读 →
论文解读

HSEmotion Team at the 11th ABAW Challenge: Multi-Task Learning and Ambivalence/Hesitancy Video Recognition

音视频 | 7.9/10

 · 更新于 2026-09-06 · 约 18 分钟 · 8619 字 阅读 →
论文解读

Hybrid Continual Learning for Low-Resource Australian Aboriginal Language Identification

语音识别 | 6.3/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7202 字 阅读 →
论文解读

Investigating the Integration of Spatial Information in Foundation-Model-Based Speaker Diarization

说话人日志 | 6.6/10

 · 更新于 2026-09-06 · 约 13 分钟 · 6032 字 阅读 →
论文解读

Listen first: Output-based multi-microphone speech enhancement

语音增强 | 6.4/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7333 字 阅读 →
论文解读

Low-Latency Neural Models for Real-Time Music Enhancement

音乐源分离 | 7.7/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6872 字 阅读 →
论文解读

Neural Morphing: Sequence-Optimized Token-Level Morphing in Neural Audio Codecs

音频编码 | 6.4/10

 · 更新于 2026-09-06 · 约 19 分钟 · 9078 字 阅读 →
论文解读

Open-Source Intelligence and Music Information Retrieval for Geographic Attribution of Musical Affect and the Ecological Limits of Population Inference

音乐理解 | 7.9/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7550 字 阅读 →
论文解读

PolarBM: Complex-valued Boltzmann Machine for Modeling Audio Signals in Polar and Log-polar Coordinates

语音增强 | 5.8/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5817 字 阅读 →
论文解读

Real-time Generation of Listener Nodding via Prediction of Kinematic Parameters for Avatar Dialogue Systems

语音交互 | 6.9/10

 · 更新于 2026-09-06 · 约 20 分钟 · 9767 字 阅读 →
论文解读

Segregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment Analysis

音频事件检测 | 6.5/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7501 字 阅读 →
论文解读

Spatial-Frequency Cued Generative Fixed-Filter Active Noise Control Based on Deep Learning in Reverberant Environments

声源定位 | 6.9/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7787 字 阅读 →
论文解读

The Sound of Absence: Audio-Language Embedding Models Struggle with Negation

音频检索 | 7.1/10

 · 更新于 2026-09-06 · 约 14 分钟 · 7003 字 阅读 →
论文解读

Traceback Translators Against Forgetting in Continual Fake Speech Detection

语音伪造检测 | 6.0/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7695 字 阅读 →
论文解读

UD-ASD: A Unified Diffusion Model for Anomalous Sound Detection

音频事件检测 | 6.6/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7397 字 阅读 →
论文解读

What is a Musical Scale? Regularity and Convention in the Organization of Pitch

音乐理解 | 5.6/10

 · 更新于 2026-09-06 · 约 12 分钟 · 5728 字 阅读 →
论文解读

ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching

语音合成 | 7.3/10

 · 更新于 2026-09-06 · 约 16 分钟 · 7717 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-07-15

共分析 25 篇语音/AI 论文

 · 更新于 2026-09-06 · 约 88 分钟 · 43611 字 阅读 →
论文解读

A Closed-Form Noise-Sensitivity Asymmetry for Causal Branch Selection in Minimal-Array TDoA Localization

音频理解 | 5.3/10

 · 更新于 2026-09-06 · 约 14 分钟 · 6698 字 阅读 →
论文解读

A Production-Oriented Framework for Evaluation of SFX Generation

音频生成 | 6.9/10

 · 更新于 2026-09-06 · 约 15 分钟 · 7493 字 阅读 →