论文解读

GLAP: General Contrastive Audio-Text Pretraining Across Domains and Languages

音频检索 | 8.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4462 字 阅读 →
论文解读

GLUE: Gradient-free Learning to Unify Experts

迁移学习 | 6.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5139 字 阅读 →
论文解读

Graph-Biased EEG Transformers for Silent Speech Decoding

语音生物标志物 | 6.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4604 字 阅读 →
论文解读

Hashing-Baseline: Rethinking Hashing in the Age of Pretrained Models

音频检索 音频分类 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4362 字 阅读 →
论文解读

Hierarchical Activity Recognition and Captioning from Long-Form Audio

音频事件检测 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4494 字 阅读 →
论文解读

High-Fidelity Speech Enhancement Via Discrete Audio Tokens

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4455 字 阅读 →
论文解读

I-DCCRN-VAE: An Improved Deep Representation Learning Framework for Complex VAE-Based Single-Channel Speech Enhancement

语音增强 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4966 字 阅读 →
会议任务专题

ICASSP 2026 - 预训练

共 1 篇 ICASSP 2026 预训练 方向论文

 · 更新于 2026-09-25 · 约 4 分钟 · 1660 字 阅读 →
论文解读

Improving Anomalous Sound Detection with Attribute-Aware Representation from Domain-Adaptive Pre-Training

音频事件检测 | 8.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5175 字 阅读 →
论文解读

Leveraging Large Speech Language Models as Evaluators for Expressive Speech

语音情感识别 | 6.5/10

 · 更新于 2026-09-25 · 约 7 分钟 · 3233 字 阅读 →
论文解读

Leveraging Multiple Speech Enhancers for Non-Intrusive Intelligibility Prediction for Hearing-Impaired Listeners

模型评估 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5540 字 阅读 →
论文解读

Leveraging Segment-Level Speech Representations for LLM-Based Speech Recognition

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5225 字 阅读 →
论文解读

Mispronunciation Detection and Diagnosis Without Model Training: A Retrieval-Based Approach

语音评估 | 8.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4234 字 阅读 →
论文解读

Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs

语音识别 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4158 字 阅读 →
论文解读

Mixture-of-Experts Based Soft-Label Learning for Multi-Label Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4604 字 阅读 →
论文解读

MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation

语音分离 | 8.0/10

 · 更新于 2026-09-25 · 约 10 分钟 · 4749 字 阅读 →
论文解读

Modeling Inter-Segment Relationships in Speech for Dementia Detection with Audio Spectrogram Transformers and Graph Attention Networks

语音生物标志物 | 7.0/10

 · 更新于 2026-09-25 · 约 9 分钟 · 4152 字 阅读 →
论文解读

MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 11 分钟 · 5146 字 阅读 →
论文解读

MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-Token Prediction

语音翻译 | 8.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5619 字 阅读 →
论文解读

Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition

语音情感识别 | 7.5/10

 · 更新于 2026-09-25 · 约 12 分钟 · 5585 字 阅读 →