论文解读

Few-Shot Open-Set Audio Classification Using Attention Information-Fused Prototypes

音频分类 | 6.8/10

 · 更新于 2026-09-25 · 约 19 分钟 · 9043 字 阅读 →
论文解读

BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations

语音识别 | 6.9/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4866 字 阅读 →
论文解读

ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models

音频分类 | 7.1/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6499 字 阅读 →
论文解读

WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation

音频分类 | 6.7/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4761 字 阅读 →
论文解读

FoleySet: A Multi-Level Human-Annotated Foley Sound Dataset

音频分类 | 7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6748 字 阅读 →
论文解读

A Multi-Stage Separation-and-Classification Framework Guided by Complementary Acoustic-to-Semantic Clues

音频分类 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4868 字 阅读 →
论文解读

Descriptor: Certus Caliber Classification Gunshot Dataset (C3GD)

音频分类 | 5.9/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4488 字 阅读 →
论文解读

Transductive Zero-Shot Audio Classification with Audio-Language Models

音频分类 | 6.4/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5272 字 阅读 →
论文解读

Turning music identification into a neural forward pass

音频分类 | 7.4/10

 · 更新于 2026-09-25 · 约 13 分钟 · 6103 字 阅读 →
论文解读

Acoustic Prompting via Stage-wise Modulation for Few-Shot Learning in Audio Language Models

音频分类 | 8.3/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5504 字 阅读 →
论文解读

MambAdapter: Lightweight Mamba-Based Adapters for Parameter-Efficient Transfer Learning in Speech and Audio

语音识别 | 8.9/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4872 字 阅读 →
论文解读

Decoding Insect Song: A Multitask Semisupervised Orthoptera Bioacoustic Classifier

音频分类 | 8.7/10

 · 更新于 2026-09-25 · 约 14 分钟 · 6708 字 阅读 →
论文解读

Dolph2Vec: Self-Supervised Representations of Dolphin Vocalizations

音频分类 | 7.2/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5440 字 阅读 →
论文解读

Few-shot Class-variable Incremental Audio Classification via Prototype Adaptation and Pseudo Class-variable Training

音频分类 | 6.9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5750 字 阅读 →
论文解读

Sound Effects Dataset Unification With the Universal Category System

音频分类 | 6.9/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5518 字 阅读 →
论文解读

Drift-Augmented Scoring: Text-Derived Noise Robustness for Zero-Shot Audio-Language Classification

音频分类 | 10/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4712 字 阅读 →
论文解读

C2GA: A Class-Controllable Generative Augmentation Framework for Respiratory Sound Classification

音频分类 | 7.3/10

 · 更新于 2026-09-25 · 约 6 分钟 · 2877 字 阅读 →
论文解读

Parameter-efficient Dual-encoder Architecture with Differentiable Choquet Integral Fusion for Underwater Acoustic Classification

音频分类 | 6.4/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5654 字 阅读 →
论文解读

几何主导下如何减出材料:GaMi 的跨模态相减式解耦

针对几何变化压制材料特征的问题,GaMi 用共位毫米波与声学的共享几何一致性做对齐-校准-相减解耦,并以样本间对比抑制残差,在 20 类材料上以 95.2% 的总体准确率显著优于单模态基线,代价是需双模态同步采集与多目标联合训练。

 · 更新于 2026-09-25 · 约 23 分钟 · 11269 字 阅读 →
论文解读

ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood

语音识别 | 8/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5167 字 阅读 →