论文解读

Does language matter for spoken word classification? A multilingual generative meta-learning approach

音频分类 | 6.0/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7080 字 阅读 →
论文解读

Better Together: Leveraging Unpaired Multimodal Data for Stronger Unimodal Models

音频分类 | 7.0/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4990 字 阅读 →
论文解读

Adaptive Embedding Fusion with Contrastive Learning for Robust Fully Few-Shot Class-Incremental Audio Classification

音频分类 | 7.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5041 字 阅读 →
论文解读

Co-Initialization of Control Filter and Secondary Path via Meta-Learning for Active Noise Control

音频安全 | 7.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5076 字 阅读 →
论文解读

DepthTalk: Few-Shot Talking Head Generation with Depth-Aware 3D Gaussian Field Motion

说话人生成 | 7.0/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3888 字 阅读 →
论文解读

Domain-Invariant Representation Learning of Bird Sounds

生物声学 | 6.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5091 字 阅读 →
论文解读

Enhancing Automatic Drum Transcription with Online Dynamic Few-Shot Learning

音乐信息检索 | 7.0/10

 · 更新于 2026-09-24 · 约 8 分钟 · 3991 字 阅读 →
论文解读

Few-Shot Recognition of Audio Deepfake Generators using Graph-Based Prototype Adaptation

音频深度伪造检测 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4546 字 阅读 →
论文解读

Lightweight and Generalizable Acoustic Scene Representations Via Contrastive Fine-Tuning and Distillation

音频场景理解 | 8.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5226 字 阅读 →
论文解读

MT-HuBERT: Self-Supervised Mix-Training for Few-Shot Keyword Spotting in Mixed Speech

关键词检测 | 7.0/10

 · 更新于 2026-09-24 · 约 23 分钟 · 11472 字 阅读 →
论文解读

PhoenixDSR: Phoneme-Guided and LLM-Enhanced Dysarthric Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5876 字 阅读 →
论文解读

Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for Voicemos 2024

语音质量评估 | 7.0/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5174 字 阅读 →
论文解读

TICL: Text-Embedding KNN for Speech in-Context Learning Unlocks Speech Recognition Abilities of Large Multimodal Models

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4710 字 阅读 →