FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
📄 FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models 标签:#音视频理解 #多模态模型 #音频理解 #Transformer #模型评估 7.9/10 | 创新 1.3/2 | 严谨 1.2/1.5 | 实验 0.8/1.5 | 清晰 1/1 | 影响 0.9/1.5 | 开源 1.5/1.5 | 复现 0.5/0.5 | 工程 0.7/1.5 ✅ 7.9/10 | 前25% | 文档类型:数据集与基准 | 评分置信度:高 | #音视频理解 | #多模态模型 | #音频理解 #Transformer | arxiv 👥 作者与机构 第一作者:Jeffrey M. Girard (Fluid Concepts Research) 通讯作者:Jeffrey M. Girard (Fluid Concepts Research, jeff@fluidconcepts.ai) 作者列表:Jeffrey M. Girard (Fluid Concepts Research), Jason Z. Zheng (Fluid Concepts Research), Jacqueline R. Vertino (Fluid Concepts Research), Antony D’Avirro (Fluid Concepts Research), Benjamin Peloquin (Fluid Concepts Research) 💡 毒舌点评 这项工作用一个“反语义泄露”的构建思路将社会感知评测推到了更可控的层面,信号检测理论的引入也巧妙揭露了模型“高分偏心”的表象。但96个dyad的小规模基准让多数模型无法显著脱离随机水平,置信区间宽如门板,削弱了排名比较的笃定性;而声称的“人类-模型统计不可区分”在如此小的样本下更像是一个统计学上的“无罪推定”而非真正的性能对齐。 ...