📄 类型未知时如何判真假:用两套耳朵听同一段音频 英文题目:Beyond Speech: Dual-Domain SSL Fusion for Unified All-Type Audio Deepfake Detection
一句话:面对语音、环境声、歌声、音乐四类混合且类型未知的真假判别,论文用 EAT-large 与 XLS-R-300M 的双域互补融合构建统一决策边界,并在评测集上以 95.58% Macro-F1 取得第二,代价是依赖封闭挑战集且未开源复现细节。
标签:#音频伪造检测 #自监督学习 #模型融合
评分:6.3/10 | 创新 1.3/2 | 技术严谨 1.2/1.5 | 实验充分 1/1.5 | 清晰度 0.8/1 | 影响力 0.9/1.5 | 开源 0/1.5 | 可复现 0.1/0.5 | 工程/实践 1/1.5
👥 作者与机构 Cunhang Fan:State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology, School of Computer Science and Technology, Anhui University, Hefei, Anhui, China Junqin Cao:State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology, School of Computer Science and Technology, Anhui University, Hefei, Anhui, China Tian Gao:Anhui Laboratory for Safe Artificial Intelligence in the Yangtze River Delta, Hefei, Anhui, China Zhipeng Xie:Anhui Laboratory for Safe Artificial Intelligence in the Yangtze River Delta, Hefei, Anhui, China Jun Xue:Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University, Wuhan, Hubei, China Zhao Lv:State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology, School of Computer Science and Technology, Anhui University, Hefei, Anhui, China Xin Fang:University of Science and Technology of China, Hefei, Anhui, China 💬 毒舌点评 用 EAT-large 管宽带声学事件、XLS-R-300M 管波形与语音细节的双域分工思路清晰,token 维拼接规避帧对齐、让注意力在统一池中自适应选线索的做法务实,保守的双重门控语音精炼也确实避免了硬路由的类型误判雪崩。但本质仍是 24 层加权求和加拼接再池化的常规融合,对伪造机理没有新假设与新约束,挑战集第二名的成绩建立在闭源、固定阈值 0.5、无显著性检验和仅在 AT-ADD 内验证的封闭环境下,可迁移性与开放域稳健性论证单薄。
...