Author Login Chief Editor Login Reviewer Login Editor Login Remote Office

Computer Engineering ›› 2026, Vol. 52 ›› Issue (9): 64-80. doi: 10.19678/j.issn.1000-3428.0253388

• Frontier Perspectives and Reviews • Previous Articles     Next Articles

Review of Multimodal False Information Detection

HAO Guanyi1, SUN Jingchao2,*()   

  1. 1. School of Investigation, People's Public Security University of China, Beijing 100038, China
    2. School of National Security, People's Public Security University of China, Beijing 100038, China
  • Received:2025-12-04 Revised:2026-03-03 Online:2026-09-15 Published:2026-04-01
  • Contact: SUN Jingchao

多模态虚假信息检测综述

郝冠一1, 孙靖超2,*()   

  1. 1. 中国人民公安大学侦查学院, 北京 100038
    2. 中国人民公安大学国家安全学院, 北京 100038
  • 通讯作者: 孙靖超
  • 作者简介:

    郝冠一(CCF学生会员), 男, 本科生, 主研方向为网络安全、侦查学

    孙靖超(通信作者), 讲师、博士

  • 基金资助:
    中国人民公安大学基本科研业务费专项资金(2023JKF02ZK04)

Abstract:

In the digital era, complex interactions among modalities such as text, image, and audio have led to multimodal misinformation. The propagation speed and concealment levels of multimodal misinformation far exceed those of traditional unimodal misinformation, posing severe challenges to information security and social governance. However, research in this field is relatively scarce in China, and a comprehensive framework has yet to be established. Therefore, this paper provides a systematic overview of the research status and development trajectory of multimodal misinformation detection. Based on a clear understanding of the core concepts and task spectrum of multimodal misinformation detection, this review provides a detailed analysis of the characteristics of the datasets and evaluation metrics. It also analyzes the applicability and detection performance of different multimodal methods and models, such as SAFE, CAFE, CFFN, SSA-MFND, PSCC-Net, DGM4, CCN, SNIFFER, and KGAlign. Three core detection methods, namely cross-modal consistency, anomaly feature recognition, and external fact-driven approaches, are summarized. Furthermore, the interpretability and generalization robustness of multimodal misinformation detection are analyzed. With the rise of Large Vision-Language Models (LVLMs), their application to multimodal misinformation detection is continuously increasing. Various application scenarios, advantages, and limitations of LVLMs in this domain are discussed. Finally, future research directions in multimodal misinformation detection are outlined, aiming to provide insights and inspiration for further development in this field.

Key words: multimodality, false information detection, Large Vision-Language Model (LVLM), modal feature fusion, fact checking

摘要:

数字时代下, 文本、图像、音频等模态的复杂交互形成了多模态虚假信息, 其传播速度与隐蔽程度远超传统单模态虚假信息, 对信息安全与社会治理构成严峻挑战。但在国内该领域相关研究较为匮乏, 尚未形成完整体系。为此, 研究系统梳理了多模态虚假信息检测领域的研究现状及发展脉络, 对多模态虚假信息检测的研究进行了全面总结。在明确多模态虚假信息检测的核心概念与任务谱系的基础上, 详细总结了数据集与测评指标特征, 分析了SAFE、CAFE、CFFN、SSA-MFND、PSCC-Net、DGM4、CCN、SNIFFER、KGAlign等不同多模态方法模型的适用场景与检测性能, 归纳了跨模态一致性、异常特征识别、外部事实驱动三大核心检测方法, 并且对多模态虚假信息检测的可解释性与泛化鲁棒性进行了探讨。同时, 随着大规模视觉语言模型(LVLM)的崛起, 其在多模态虚假信息检测中的应用不断深化, 对此研究梳理了LVLM在该领域的多种应用场景、优势与局限。最后展望了多模态虚假信息检测的未来研究方向, 以期为多模态虚假信息检测领域的发展提供借鉴与启示。

关键词: 多模态, 虚假信息检测, 大规模视觉语言模型, 模态特征融合, 事实核查