作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (7): 56-75. doi: 10.19678/j.issn.1000-3428.0070621

• 前沿观点与综述 • 上一篇    下一篇

面向网络威胁情报领域的信息抽取综述

冯嘉琦, 高见*()   

  1. 中国人民公安大学信息网络安全学院, 北京 100038
  • 收稿日期:2024-11-18 修回日期:2024-12-28 出版日期:2026-07-15 发布日期:2026-07-04
  • 通讯作者: 高见
  • 作者简介:

    冯嘉琦, 男, 硕士研究生, 主研方向为网络安全威胁情报、自然语言处理

    高见(通信作者), 副教授、博士

  • 基金资助:
    中国人民公安大学中央高校基本科研业务费专项资金(2024JKF17)

Review of Information Extraction in the Field of Cyber Threat Intelligence

FENG Jiaqi, GAO Jian*()   

  1. School of Information and Network Security, People's Public Security University of China, Beijing 100038, China
  • Received:2024-11-18 Revised:2024-12-28 Online:2026-07-15 Published:2026-07-04
  • Contact: GAO Jian

摘要:

网络威胁情报(CTI)通过分析和预测网络威胁, 为网络安全防护体系的建设提供重要支持。信息抽取(IE)技术作为CTI领域的核心基础, 近年来在学术界和工业界受到了广泛关注, 尤其在威胁情报知识图谱的构建中, 其应用至关重要。面对不断演变的威胁类型和攻击手段, 精确的信息抽取能力对于提升网络安全防御水平具有决定性的影响。本研究首先系统性地回顾了CTI领域信息抽取技术的发展, 详细梳理并总结了信息抽取框架的关键构成要素; 然后将信息抽取方法归类为基于传统机器学习的方法、基于深度学习的方法以及基于大语言模型(LLM)的方法三大类别, 并对每类方法的特点、适用场景及研究进展进行了详尽的阐述; 最后对现有技术进行了定量分析和理论总结, 深入探讨了其在实际应用中的创新点和局限性, 同时梳理了当前领域面临的挑战, 并针对这些挑战提出了潜在的研究方向。通过对CTI信息抽取技术的全面综述, 本研究为该领域研究提供了一个全面的学术视角, 并为促进更高效、更智能的网络安全防护体系的发展提供了重要的理论基础。

关键词: 网络安全, 网络威胁情报, 信息抽取, 机器学习, 深度学习, 大语言模型

Abstract:

Cyber Threat Intelligence (CTI) provides critical support for the construction of cybersecurity defense systems by analyzing and predicting cyber threats. Information Extraction (IE) technology, which is a foundational component of CTI, has garnered significant attention from both academia and industry in recent years. Its application is particularly important for constructing threat intelligence knowledge graphs. Given the ever-evolving nature of threat types and attack techniques, precise IE capabilities play a decisive role in enhancing cybersecurity defense. This review systematically examines the development of IE technologies within the CTI domain and presents a comprehensive analysis and summary of the key components of IE frameworks. IE methods are categorized into three major types: traditional machine-learning-based, deep-learning-based, and large language-model-based methods. The characteristics, applicable scenarios, and progress of research in each category are discussed in detail. Moreover, quantitative and theoretical analyses of existing techniques are presented, along with a discussion of their implications and limitations in practical applications. The review also outlines the current challenges in the field and proposes potential research directions to address them. By presenting a detailed assessment of IE technologies for CTI, this review offers a comprehensive academic perspective and lays an important theoretical foundation for fostering more efficient and intelligent cybersecurity defense systems.

Key words: cybersecurity, Cyber Threat Intelligence (CTI), Information Extraction (IE), machine learning, deep learning, Large Language Model (LLM)