论文解读

TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios

音频分类 | 7.4/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6741 字 阅读 →
论文解读

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types

音频分类 | 6.1/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7376 字 阅读 →
论文解读

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition

语音识别 | 6.5/10

 · 更新于 2026-09-24 · 约 14 分钟 · 7010 字 阅读 →
论文解读

SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation

音频伪造检测 | 6.5/10

 · 更新于 2026-09-24 · 约 19 分钟 · 9156 字 阅读 →
论文解读

Convex Low-resource Accent-Robust Language Detection in Speech Recognition

语音识别 | 6/10

 · 更新于 2026-09-24 · 约 15 分钟 · 7467 字 阅读 →
论文解读

Enhancing Acoustic-to-Articulatory Inversion with Multi-Target Pretraining for Low-Resource Settings

语音交互 | 7/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5177 字 阅读 →
论文解读

Disentangling Speaker and Language Effects in Cross-Lingual Speaker Verification for Iberian Languages

说话人验证 | 5.6/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5518 字 阅读 →
论文解读

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models

音频分类 | 7.6/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4548 字 阅读 →
论文解读

Two kinds of robustness are not the same: disentangling fault tolerance and low-SNR robustness in multi-domain event detection on real data

音频事件检测 | 8.9/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4694 字 阅读 →
论文解读

wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2

自监督学习 | 8.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5028 字 阅读 →
论文解读

Do Speech Emphasis Models Generalize across Languages and Emotions?

语音识别 | 7/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4935 字 阅读 →
论文解读

FoleySet: A Multi-Level Human-Annotated Foley Sound Dataset

音频分类 | 7/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6748 字 阅读 →
论文解读

Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR

语音识别 | 9/10

 · 更新于 2026-09-24 · 约 14 分钟 · 6751 字 阅读 →
论文解读

DSSCNet: A Transfer Learning Framework for Cross-Corpus Dysarthric Speech Severity Classification

迁移学习 | 6.3/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5528 字 阅读 →
论文解读

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era

语音识别 | 7.5/10

 · 更新于 2026-09-24 · 约 23 分钟 · 11361 字 阅读 →
论文解读

How Well Do Self-Supervised Speech Models Encode Age and Gender in Children's Speech? A Layer-Wise Analysis Across Multiple Architectures

自监督学习 | 9/10

 · 更新于 2026-09-24 · 约 10 分钟 · 4728 字 阅读 →
论文解读

Leveraging systems' non-linearity to tackle the scarcity of data in the design of Intelligent Fault Diagnosis Systems

数据增强 | 5.5/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5424 字 阅读 →
论文解读

Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning

语音识别 | 8.7/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5408 字 阅读 →
论文解读

Low-resource Language Discrimination Towards Chinese Dialects with Transfer learning and Data Augmentation

语音识别 | 5.5/10

 · 更新于 2026-09-24 · 约 12 分钟 · 5615 字 阅读 →
论文解读

Perceptual compensation for tonal context in self-supervised speech models

语音识别 | 7.7/10

 · 更新于 2026-09-24 · 约 11 分钟 · 5350 字 阅读 →