A Unified Uncertainty-Aware Back-End for Speaker Verification: Scoring, Normalization, and Calibration

📄 当嵌入不再确定:把不确定性从打分一路带到校准的后端闭环 英文题目:A Unified Uncertainty-Aware Back-End for Speaker Verification: Scoring, Normalization, and Calibration 一句话:针对试次可靠性随长度与信道而变的说话人验证,论文用同一对角协方差贯穿余弦打分、队列归一化与逻辑回归校准,在 VoxCeleb 双骨干上将平均相对提升推至约 28 % 与 27 %,代价是增益高度依赖不确定性估计质量且部分 minDCF 出现回退。 标签:#说话人验证 | #对比学习 | #鲁棒性 | #模型融合 | #基准测试 评分:7.8/10 | 创新 1.3/2 | 技术严谨 1.2/1.5 | 实验充分 1/1.5 | 清晰度 0.8/1 | 影响力 1/1.5 | 开源 1/1.5 | 可复现 0.3/0.5 | 工程/实践 1.2/1.5 👥 作者与机构 Junjie Li:机构信息未在 arXiv HTML 中可靠披露 Kong Aik Lee:机构信息未在 arXiv HTML 中可靠披露 💬 毒舌点评 亮点在于把协方差从打分一路用到归一化和校准,补上了以往不确定性只停留在前端的断层,工程闭环做得完整且在双骨干上均有提升。短板是三段式增益拆开看每个增量都很小且部分 minDCF 反而变差,过度依赖 U^3-xi 这一特定不确定性估计器,对不确定性质量本身缺乏鲁棒性分析。 ...

2026-09-02 · 更新于 2026-09-04 · 5 min · 1064 words

Beyond Speech: Dual-Domain SSL Fusion for Unified All-Type Audio Deepfake Detection

📄 类型未知时如何判真假:用两套耳朵听同一段音频 英文题目:Beyond Speech: Dual-Domain SSL Fusion for Unified All-Type Audio Deepfake Detection 一句话:面对语音、环境声、歌声、音乐四类混合且类型未知的真假判别,论文用 EAT-large 与 XLS-R-300M 的双域互补融合构建统一决策边界,并在评测集上以 95.58% Macro-F1 取得第二,代价是依赖封闭挑战集且未开源复现细节。 标签:#音频伪造检测 #自监督学习 #模型融合 评分:6.3/10 | 创新 1.3/2 | 技术严谨 1.2/1.5 | 实验充分 1/1.5 | 清晰度 0.8/1 | 影响力 0.9/1.5 | 开源 0/1.5 | 可复现 0.1/0.5 | 工程/实践 1/1.5 👥 作者与机构 Cunhang Fan:State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology, School of Computer Science and Technology, Anhui University, Hefei, Anhui, China Junqin Cao:State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology, School of Computer Science and Technology, Anhui University, Hefei, Anhui, China Tian Gao:Anhui Laboratory for Safe Artificial Intelligence in the Yangtze River Delta, Hefei, Anhui, China Zhipeng Xie:Anhui Laboratory for Safe Artificial Intelligence in the Yangtze River Delta, Hefei, Anhui, China Jun Xue:Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University, Wuhan, Hubei, China Zhao Lv:State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology, School of Computer Science and Technology, Anhui University, Hefei, Anhui, China Xin Fang:University of Science and Technology of China, Hefei, Anhui, China 💬 毒舌点评 用 EAT-large 管宽带声学事件、XLS-R-300M 管波形与语音细节的双域分工思路清晰,token 维拼接规避帧对齐、让注意力在统一池中自适应选线索的做法务实,保守的双重门控语音精炼也确实避免了硬路由的类型误判雪崩。但本质仍是 24 层加权求和加拼接再池化的常规融合,对伪造机理没有新假设与新约束,挑战集第二名的成绩建立在闭源、固定阈值 0.5、无显著性检验和仅在 AT-ADD 内验证的封闭环境下,可迁移性与开放域稳健性论证单薄。 ...

2026-09-01 · 更新于 2026-09-04 · 7 min · 1461 words

Investigating voiced and unvoiced regions of speech for audio deepfake detection

📄 Investigating voiced and unvoiced regions of speech for audio deepfake detection 标签:#语音伪造检测 #图神经网络 #模型融合 #模型评估 6.4/10 | 创新 1.2/2 | 严谨 1.1/1.5 | 实验 1.2/1.5 | 清晰 0.9/1 | 影响 1/1.5 | 开源 0/1.5 | 复现 0.3/0.5 | 工程 0.7/1.5 ✅ 6.4/10 | 前50% | 文档类型:方法研究 | 评分置信度:中 | #语音伪造检测 | #图神经网络 | #模型融合 #模型评估 | arxiv 👥 作者与机构 第一作者:Ganesh Sivaraman(Pindrop, Atlanta, USA) 通讯作者:正文未明确标注通讯作者 作者列表:Ganesh Sivaraman、Hemlata Tak、Elie Khoury(机构:Pindrop, Atlanta, USA) ...

2026-08-26 · 更新于 2026-09-04 · 5 min · 994 words

SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception

📄 SonarLLM: A Native Sonar–Optical Multimodal Large Language Model for Underwater Perception 标签:#音频理解 #多模态模型 #模型融合 #鲁棒性 8.0/10 | 创新 1.7/2 | 严谨 1.4/1.5 | 实验 1.4/1.5 | 清晰 0.9/1 | 影响 1.1/1.5 | 开源 0/1.5 | 复现 0.4/0.5 | 工程 1.1/1.5 🔥 8.0/10 | 前25% | 文档类型:方法研究 | 评分置信度:高 | #音频理解 | #多模态模型 | #模型融合 #鲁棒性 | arxiv 👥 作者与机构 第一作者:Cong Su(Faculty of Information Engineering and Automation, Kunming University of Science and Technology;Yunnan Key Laboratory of Artificial Intelligence) 通讯作者:Longxuan Ma 作者列表:Cong Su、Longxuan Ma、Ling Dong、Guofeng Tang、Weijie Yin、Haohui Chen、Zhengtao Yu(机构:Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming, China;Yunnan Key Laboratory of Artificial Intelligence, Kunming, China) ...

2026-08-26 · 更新于 2026-09-04 · 6 min · 1102 words

Adaptive Hierarchical Representation Alliance for Multimodal Learning

📄 Adaptive Hierarchical Representation Alliance for Multimodal Learning 标签:#语音情感识别 #多模态模型 #模型融合 #鲁棒性 8.6/10 | 创新 1.7/2 | 严谨 1.4/1.5 | 实验 1.4/1.5 | 清晰 0.9/1 | 影响 1.4/1.5 | 开源 0.2/1.5 | 复现 0.4/0.5 | 工程 1.2/1.5 🔥 8.6/10 | 前25% | 文档类型:方法研究 | 评分置信度:中 | #语音情感识别 | #多模态模型 | #模型融合 #鲁棒性 | arxiv 👥 作者与机构 第一作者:Chunlei Meng(College of Intelligent Robotics and Advanced Manufacturing, Fudan University) 通讯作者:Chunlei Meng、Chun Ouyang 作者列表:Chunlei Meng、Pengbin Feng、Jacqueline J. Pang、Chih-Ting Liao、Rong Fu、Zhaolu Kang、Zhongxue Gan、Chun Ouyang(机构:Fudan University;University of Southern California;J.P. Morgan Chase;University of New South Wales;Independent Researcher;Peking University) ...

2026-08-25 · 更新于 2026-09-04 · 2 min · 399 words

Mitigating Speaker Leakage in Cascaded Multi-talker ASR with Diarization-based Transcript Correction

📄 Mitigating Speaker Leakage in Cascaded Multi-talker ASR with Diarization-based Transcript Correction 标签:#语音识别 #模型融合 #语音分离 #说话人日志 #会议转录 7.1/10 | 创新 1.5/2 | 严谨 1.3/1.5 | 实验 1.3/1.5 | 清晰 0.8/1 | 影响 1.1/1.5 | 开源 0/1.5 | 复现 0.3/0.5 | 工程 0.8/1.5 ✅ 7.1/10 | 前50% | 文档类型:方法研究 | 评分置信度:中 | #语音识别 | #模型融合 | #语音分离 #说话人日志 | arxiv 👥 作者与机构 第一作者:Hermann Yepdjio Nkouanga(Portland State University, USA) 通讯作者:正文未明确标注 作者列表:Hermann Yepdjio Nkouanga、Minwei Luo、Maggie Wigness、Suresh Singh(机构:Portland State University, USA;US Army Research Laboratory, USA) ...

2026-08-25 · 更新于 2026-09-04 · 3 min · 541 words

Multi-Modal Semantic Expansion with Constrained LLM Reranking for Conversational Music Recommendation

📄 Multi-Modal Semantic Expansion with Constrained LLM Reranking for Conversational Music Recommendation 标签:#音乐推荐 #模型融合 #大语言模型 #音乐文本检索 8.3/10 | 创新 1.5/2 | 严谨 1.3/1.5 | 实验 1.4/1.5 | 清晰 0.9/1 | 影响 1.4/1.5 | 开源 0.2/1.5 | 复现 0.5/0.5 | 工程 1.1/1.5 🔥 8.3/10 | 前25% | 文档类型:系统技术报告 | 评分置信度:中 | #音乐推荐 | #模型融合 | #大语言模型 #音乐文本检索 | arxiv 👥 作者与机构 第一作者:Naman Garg(National Institute of Technology Kurukshetra, India) 通讯作者:正文未明确标注 作者列表:Naman Garg、Sarika Jain、George Fazekas(机构:National Institute of Technology Kurukshetra, India;Queen Mary University of London, United Kingdom) ...

2026-08-25 · 更新于 2026-09-04 · 3 min · 563 words

TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems

📄 TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems 标签:#语音识别 #模型融合 #流式处理 #实时处理 #高效推理 9.0/10 | 创新 1.6/2 | 严谨 1.3/1.5 | 实验 1.4/1.5 | 清晰 0.8/1 | 影响 1.2/1.5 | 开源 1.2/1.5 | 复现 0.4/0.5 | 工程 1.1/1.5 🔥 9.0/10 | 前10% | 文档类型:系统技术报告 | 评分置信度:中 | #语音识别 | #模型融合 | #流式处理 #实时处理 | arxiv 👥 作者与机构 第一作者:Vladimir Bataev(NVIDIA, Yerevan, Armenia) 通讯作者:Vladimir Bataev 作者列表:Vladimir Bataev、Lilit Grigoryan、Andrei Andrusenko、Nikolay Karpov、Vitaly Lavrukhin、Boris Ginsburg(机构:NVIDIA, Yerevan, Armenia;NVIDIA, Santa Clara, USA) ...

2026-08-25 · 更新于 2026-09-04 · 4 min · 711 words

Multimodal Rapport Estimation in Real-World HRI

📄 Multimodal Rapport Estimation in Real-World HRI 标签:#多模态模型 #音视频理解 #模型融合 #模型评估 7.0/10 | 创新 1.1/2 | 严谨 1.1/1.5 | 实验 1.2/1.5 | 清晰 0.9/1 | 影响 0.9/1.5 | 开源 0.5/1.5 | 复现 0.3/0.5 | 工程 1/1.5 ✅ 7.0/10 | 前50% | 文档类型:应用研究 | 评分置信度:中 | #多模态模型 | #音视频理解 | #模型融合 #模型评估 | arxiv 👥 作者与机构 第一作者:Akihiro Sakuramoto(机构未说明) 通讯作者:未说明 作者列表:Akihiro Sakuramoto、Takato Hayashi、Ryo Miyoshi、Yuki Okafuji、Shogo Okada(机构信息未在当前正文中完整说明) 💡 毒舌点评 真实药店环境的人机交互亲和力估计是少见的实地研究,但要把结果当作 HRI 结论需要过四道折扣:单一日本药店、单一机器人形态、日语文化环境,而且机器人由真人远程操控——自主交互的 rapport 动态可能很不一样。九十七个个体样本经分层后更小,三人条件只剩十三个,趋势可能由少数 session 驱动。算法估计的是第三方观察到的亲和度而非参与者自评,不应被读作内心状态测量。低语音与极早退出的互动恰好是被过滤掉的、真实部署中最难的部分。 ...

2026-08-20 · 更新于 2026-09-04 · 3 min · 596 words

AT-ADD: All-Type Audio Deepfake Detection Challenge Summary

📄 AT-ADD: All-Type Audio Deepfake Detection Challenge Summary 标签:#音频伪造检测 #自监督学习 #语音伪造检测 #模型融合 #基准测试 6.3/10 | 创新 1.2/2 | 严谨 1/1.5 | 实验 0.7/1.5 | 清晰 0.8/1 | 影响 1/1.5 | 开源 0.5/1.5 | 复现 0.1/0.5 | 工程 1/1.5 ✅ 6.3/10 | 前50% | 文档类型:数据集与基准 | 评分置信度:中 | #音频伪造检测 | #自监督学习 | #语音伪造检测 #模型融合 | arxiv 👥 作者与机构 第一作者:Yuankun Xie(Communication University of China & Ant Group,北京,中国;原文标注共同贡献) 通讯作者:未说明 作者列表:Yuankun Xie(Communication University of China & Ant Group)、Haonan Cheng(Communication University of China)、Jiayi Zhou(Machine Intelligence, Ant Group,上海)、Xiaoxuan Guo(Communication University of China & Ant Group)、Tao Wang(Machine Intelligence, Ant Group,上海)、Changhao Zhang(Machine Intelligence, Ant Group,上海)、Jian Liu(Machine Intelligence, Ant Group,上海)、Weiqiang Wang(Machine Intelligence, Ant Group,上海)、Ruibo Fu(Institute of Automation, Chinese Academy of Sciences,北京)、Xiaopeng Wang(Beijing Institute of Technology,北京)、Hengyan Huang(Communication University of China)、Xiaoying Huang(Communication University of China)、Long Ye(Communication University of China)、Guangtao Zhai(Shanghai Jiao Tong University,上海) 💡 毒舌点评 该工作总结了覆盖面较广的音频伪造检测挑战,数据规模和双赛道设计有实际价值,尤其是将语音、环境声、歌声和音乐纳入统一评测。短板也很明显:评估集未见生成器只给数量、不给清单,缺少按类型/生成器/扰动的细粒度分析,参赛系统描述更像榜单笔记而非可复现技术报告。论文没有官方基线、消融或统计显著性检验,因此“all-type”和“robust”的性能声明更多停留在排行榜数字层面。 ...

2026-08-17 · 更新于 2026-09-04 · 4 min · 821 words