CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model
📄 CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model 标签:#语音质量评估 #多模态模型 #大语言模型 #参数高效微调 #教育 7.9/10 | 创新 1.3/2 | 严谨 1.2/1.5 | 实验 1.1/1.5 | 清晰 0.9/1 | 影响 0.9/1.5 | 开源 1/1.5 | 复现 0.3/0.5 | 工程 1.2/1.5 ✅ 7.9/10 | 前25% | 文档类型:方法研究 | 评分置信度:高 | #语音质量评估 | #多模态模型 | #大语言模型 #参数高效微调 | arxiv 👥 作者与机构 第一作者:Nhan Phan(Aalto University, Department of Information and Communications Engineering) 通讯作者:未说明 作者列表: Nhan Phan(Aalto University, Department of Information and Communications Engineering) Ilona Lähteenmäki(University of Helsinki, Department of Education) Anna von Zansen(University of Helsinki, Department of Education) Olli-Pekka Pauna(Aalto University, Department of Information and Communications Engineering) Yaroslav Getman(Aalto University, Department of Information and Communications Engineering) Tamás Grósz(Aalto University, Department of Information and Communications Engineering / Walton Institute, Programmable Autonomous Systems Division) Mikko Kurimo(Aalto University, Department of Information and Communications Engineering) 💡 毒舌点评 该工作的双分支设计把"声学传达"与"文本内容"干净地拆开,并用带容差的辅助损失和 LoRA 适配 Whisper 中间层,提供了比大多数 speech-LLM 系统更少参数的可解释方案,实验透明度较好。然而,其 SOTA 声明建立在 RMSE 相差 0.002 的边际优势上,作者自己都不承认有实质改进,且没有跨数据集验证,距离真正的可泛化口语评估系统还有相当距离。 ...