论文解读

VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track

音频问答 | 3.9/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5843 字 阅读 →
论文解读

DAStatFormer: A Hybrid Multibranch Transformer with Statistical Feature Integration for DAS-Based Pattern Recognitions

音频事件检测 | 6.4/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4723 字 阅读 →
论文解读

Improving acoustic drone detection generalization through pretraining and data augmentation

音频事件检测 | 7.7/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5246 字 阅读 →
论文解读

Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems

语音识别 | 6.8/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6086 字 阅读 →
论文解读

Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5492 字 阅读 →
论文解读

A conceptual framework for learning to listen by reward: Curiosity-driven search for novel sources

声源定位 | 4.0/10

 · 更新于 2026-09-24 · 约 18 分钟 · 8794 字 阅读 →
论文解读

A strongly annotated passive acoustic dataset for tropical bird monitoring

生物声学 | 7.2/10

 · 更新于 2026-09-24 · 约 21 分钟 · 10188 字 阅读 →
论文解读

Executable Boundary Contracts for Sound Event Traces

音频事件检测 | 8.5/10

 · 更新于 2026-09-24 · 约 20 分钟 · 9595 字 阅读 →
论文解读

SEABAD: A Tropical Bird Activity Detection Dataset for Passive Acoustic Monitoring

生物声学 音频事件检测 | 8.1/10

 · 更新于 2026-09-24 · 约 17 分钟 · 8250 字 阅读 →
论文解读

Executable Boundary Contracts for Sound Event Traces

音频事件检测 | 8.4/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9164 字 阅读 →
论文解读

AudioMosaic: Contrastive Masked Audio Representation Learning

音频分类 | 7.3/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7392 字 阅读 →
论文解读

FSD50K-Solo: Automated Curation of Single-Source Sound Events

数据清洗 | 5.5/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6839 字 阅读 →
论文解读

Physics-Based iOCT Sonification for Real-time Interaction Awareness in Subretinal Injection

医疗音频 | 6.5/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7411 字 阅读 →
论文解读

NAACA: Training-Free NeuroAuditory Attentive Cognitive Architecture with Oscillatory Working Memory for Salience-Driven Attention Gating

音频事件检测 | 7.0/10

 · 更新于 2026-09-24 · 约 16 分钟 · 7828 字 阅读 →
论文解读

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing

音频事件检测 | 5.8/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9380 字 阅读 →
论文解读

BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing

BeeVe: Unsupervised Acoustic State Discovery in Honey Bee Buzzing

 · 更新于 2026-09-24 · 约 20 分钟 · 9725 字 阅读 →
论文解读

MultiLinguahah : A New Unsupervised Multilingual Acoustic Laughter Segmentation Method

音频事件检测 | 8.5/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6396 字 阅读 →
论文解读

Towards Open World Sound Event Detection

音频事件检测 | 8.5/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6369 字 阅读 →
论文解读

Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning

音视频 | 7.5/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6444 字 阅读 →
论文解读

HARMES: A Multi-Modal Dataset for Wearable Human Activity Recognition with Motion, Environmental Sensing and Sound

音频分类 | 8.0/10

 · 更新于 2026-09-24 · 约 13 分钟 · 6251 字 阅读 →