论文解读

Profy: Interpretable Visualization of Expertise-Dependent Motor Skills Toward Supporting Piano Practice

音乐信息检索 | 6.9/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5822 字 阅读 →
论文解读

RAT: Reference-Augmented Training for ASV Anti-Spoofing

数据增强 | 8.8/10

 · 更新于 2026-09-08 · 约 10 分钟 · 4995 字 阅读 →
论文解读

Recovering the Zipfian Distribution in Unsupervised Term Discovery

自监督学习 | 8.7/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5173 字 阅读 →
论文解读

RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification

对比学习 | 6.5/10

 · 更新于 2026-09-08 · 约 13 分钟 · 6068 字 阅读 →
论文解读

Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding

多模态模型 | 9.4/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5947 字 阅读 →
论文解读

Speaker Group Encoding in Self-supervised Speech Recognition Models

语音识别 | 6.5/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5916 字 阅读 →
论文解读

Speech Encoder Fusion for LLM-based Automatic Speech Recognition

语音识别 | 7.2/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5281 字 阅读 →
论文解读

Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation

语音识别 | 8.3/10

 · 更新于 2026-09-08 · 约 13 分钟 · 6269 字 阅读 →
论文解读

SSL-GMMVC: Interpretable Voice Conversion via Locally Linear GMM Transforms in Self-Supervised Representation Space

语音转换 | 6.8/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5852 字 阅读 →
论文解读

Time-frequency localization of bird calls in dense soundscapes

信号处理基础 | 8.5/10

 · 更新于 2026-09-08 · 约 13 分钟 · 6377 字 阅读 →
论文解读

Towards Deep Contextual Reasoning from Broad Descriptions for ASR with Speech-LLM via Metadata-Driven Reasoning Chains

语音识别 | 6.2/10

 · 更新于 2026-09-08 · 约 10 分钟 · 4631 字 阅读 →
论文解读

Towards Robust Arabic Speech Emotion Recognition with Deep Learning

语音情感识别 | 6.4/10

 · 更新于 2026-09-08 · 约 12 分钟 · 5530 字 阅读 →
论文解读

TRADE: Transducer-Augmented Decoder for Speech LLM

语音识别 | 7.4/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5157 字 阅读 →
论文解读

ViP-VL: Vietnamese Self-supervised Speech Pretraining Model with Vector-Quantization Learning

语音识别 | 9.7/10

 · 更新于 2026-09-08 · 约 11 分钟 · 5467 字 阅读 →
论文解读

What Do Deepfake Speech Detectors Actually Hear?

What Do Deepfake Speech Detectors Actually Hear?

 · 更新于 2026-09-08 · 约 2 分钟 · 644 字 阅读 →
每日研究速递

语音/音乐/音频论文速递 2026-06-10

共分析 45 篇语音/AI 论文

 · 更新于 2026-09-08 · 约 126 分钟 · 62690 字 阅读 →
论文解读

A Comparative Study of Pre-trained Speech Encoders and Training Objectives for Large-Scale Indic Spoken Language Identification

自监督学习 | 8.9/10

 · 更新于 2026-09-08 · 约 10 分钟 · 4667 字 阅读 →
论文解读

A Comparison of SSL-Based Feature Extractors and Back-End Classifiers for Spoofing Detection: A Multi-Corpus Training and Cross-Linguistic Analysis

自监督学习 | 5/10

 · 更新于 2026-09-08 · 约 22 分钟 · 10637 字 阅读 →
论文解读

A Finetuned SpeechLLM for Joint Multi-Granular L2 Assessment and Natural-Language Rationales

大语言模型 | 10/10

 · 更新于 2026-09-08 · 约 16 分钟 · 7516 字 阅读 →
论文解读

A Hierarchical Feature Engineering Framework for Automated Classification of Phonotraumatic and Non-Phonotraumatic Vocal Hyperfunction

A Hierarchical Feature Engineering Framework for Automated Classification of Phonotraumatic and Non-Phonotraumatic Vocal Hyperfunction

 · 更新于 2026-09-08 · 约 12 分钟 · 5753 字 阅读 →