论文解读
Do speech foundation models perceive speaker similarity as humans do?
说话人识别 | 6.3/10
说话人识别 | 6.3/10
图神经网络 | 6.6/10
大语言模型 | 6.3/10
Enhancing Audio Captioning with Auxiliary AudioSet Semantics
音乐生成 | 7.7/10
语音合成 | 7.2/10
语音识别 | 8.0/10
音频生成 | 7.5/10
音频检索 | 7.4/10
参数高效微调 | 8.1/10
语音合成 | 8.2/10
InfoShield: Privacy-Preserving Speech Representations for Mental Health Screening via Information-Theoretic Optimization
语音情感识别 | 8.1/10
语音识别 | 9/10
语音识别 | 8.4/10
语音识别 | 6.9/10
迁移学习 | 5.7/10
开源工具 | 7.5/10
语音翻译 | 8.6/10
Probing Spatial Structure in Pretrained Audio Representations