论文解读

TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation

语音分离 | 8.1/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5896 字 阅读 →
论文解读

TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech

TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech

 · 更新于 2026-09-07 · 约 12 分钟 · 5670 字 阅读 →
论文解读

Two kinds of robustness are not the same: disentangling fault tolerance and low-SNR robustness in multi-domain event detection on real data

音频事件检测 | 8.9/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4694 字 阅读 →
论文解读

Underwater Source Detection and Classification for Signal-based Surveillance: Audio Dataset Curation and Cross-Domain Evaluation

数据集 | 7.8/10

 · 更新于 2026-09-07 · 约 9 分钟 · 4494 字 阅读 →
论文解读

VeRe-Flow: Guiding Flow Matching toward Clean Speech via Velocity Contrastive Regularization and Representation Alignment for Noise-Robust Bandwidth Expansion

语音增强 | 7.7/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4813 字 阅读 →
论文解读

VIB-AVSR: Variational Information Bottleneck for Noise-Robust LLM-Based Audio-Visual Speech Recognition

语音识别 | 9/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5609 字 阅读 →
论文解读

wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2

自监督学习 | 8.5/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5028 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-30

共分析 35 篇语音/AI 论文

 · 更新于 2026-09-07 · 约 102 分钟 · 51078 字 阅读 →
论文解读

A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset

A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset

 · 更新于 2026-09-07 · 约 13 分钟 · 6342 字 阅读 →
论文解读

A Survey of Automated Presentation Coaching: Systems, Methods, and Open Challenges

语音识别 | 5.4/10

 · 更新于 2026-09-07 · 约 13 分钟 · 6247 字 阅读 →
论文解读

Advancing Speaker-Based Vocal Effort Classification with WavLM and Data Augmentation in Naturalistic Non-Calibrated Speech Recordings

语音增强 | 6.8/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4639 字 阅读 →
论文解读

DG^VoiC: Speaker Clustering for Fraud Investigation under Real Call-Centre Conditions

说话人识别 | 5.7/10

 · 更新于 2026-09-07 · 约 9 分钟 · 4463 字 阅读 →
论文解读

Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection

说话人识别 | 6/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5549 字 阅读 →
论文解读

Do Speech Emphasis Models Generalize across Languages and Emotions?

语音识别 | 7/10

 · 更新于 2026-09-07 · 约 10 分钟 · 4935 字 阅读 →
论文解读

From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection

语音识别 | 6/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5871 字 阅读 →
论文解读

From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection

数据增强 | 8.3/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5898 字 阅读 →
论文解读

Grammar-Guided Hierarchical Parsing for Long-form Audio Activity Recognition

音频事件检测 | 6.2/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5140 字 阅读 →
论文解读

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

语音合成 | 8.2/10

 · 更新于 2026-09-07 · 约 27 分钟 · 13336 字 阅读 →
论文解读

HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models

语音合成 | 6.5/10

 · 更新于 2026-09-07 · 约 12 分钟 · 5997 字 阅读 →
论文解读

Learning from Annotation Uncertainty: Entropy-Aware Curriculum for Speech Emotion Recognition

语音情感识别 | 7.4/10

 · 更新于 2026-09-07 · 约 11 分钟 · 5469 字 阅读 →