作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (10): 363-374. doi: 10.19678/j.issn.1000-3428.0070706

• 大模型与生成式人工智能 • 上一篇    

基于大模型与强化学习的网络威胁狩猎方法

崔泽源1, 葛文翰2, 王俊峰1,2   

  1. 1. 四川大学视觉合成图形图像技术国防重点学科实验室, 四川 成都 610065;
    2. 四川大学计算机学院, 四川 成都 610065
  • 收稿日期:2024-12-16 修回日期:2025-03-20 发布日期:2025-05-08
  • 作者简介:崔泽源,男,硕士研究生,主研方向为网络与信息安全;葛文翰,博士研究生;王俊峰(通信作者),教授、博士,E-mail:wangjf@scu.edu.cn。
  • 基金资助:
    国家自然科学基金(U24B20147,U2133208);四川省重点研发计划(2024ZHCG0195,2024ZDZX0044,2024ZYD0269)。

Cyber Threat Hunting Method Based on Large Language Model and Reinforcement Learning

CUI Zeyuan1, GE Wenhan2, WANG Junfeng1,2   

  1. 1. National Key Laboratory of Fundamental Science on Synthetic Vision, Sichuan University, Chengdu 610065, Sichuan, China;
    2. College of Computer Science, Sichuan University, Chengdu 610065, Sichuan, China
  • Received:2024-12-16 Revised:2025-03-20 Published:2025-05-08

摘要: 网络威胁狩猎(CTH)通过主动发现攻击线索与恶意证据实现对攻击事件的快速响应。现有网络威胁狩猎方法虽具备在广泛信息源条件下进行决策的能力,但在现实场景中存在先验知识缺乏以及反馈稀疏的问题。针对上述问题,提出一种基于大语言模型(LLM)和强化学习(RL)的网络威胁狩猎算法RE-HUNTER。为解决先验知识缺乏的问题,该方法构建了上下文向量数据库,利用大语言模型中的领域知识和网络威胁情报中的非结构化知识提升决策方法的冷启动效果,初始化强化学习权重;为解决反馈稀疏的问题,该方法改进了蒙特卡洛树搜索(MCTS)算法,引入了递归更新机制和方法相似度机制,从而增强了对实体和方法执行结果的反馈。基于186个真实攻击案例进行的网络威胁狩猎实验结果显示,该方法相较当前最优基线方法显著提高了搜索效率,在0~2 000步区间内平均召回率相对提升18.24%;值得注意的是,在0~250步冷启动区间内较基线方法实现了86.28%的平均召回率提升;此外,消融实验表明该方法的不同组成部分对实验结果均起到正向作用,能够有效降低网络威胁狩猎的成本。

关键词: 网络威胁狩猎, 主动防御, 强化学习, 大语言模型, 蒙特卡洛树搜索

Abstract: Cyber Threat Hunting (CTH) enables rapid response to attack events through proactive discovery of attack clues and malicious evidence. Although existing CTH algorithms can search extensive information sources, they face challenges in real-world scenarios owing to insufficient prior knowledge and sparse feedback. To address these problems, this paper proposes a CTH algorithm, RE-HUNTER, based on Large Language Model (LLM) and Reinforcement Learning (RL). To address the lack of prior knowledge, this algorithm constructs a contextual vector database and leverages domain expertise from LLM and unstructured knowledge from cyber threat intelligence for cold-start decision-making to initialize RL weights. To address sparse feedback, this algorithm improves the Monte Carlo Tree Search (MCTS) algorithm by introducing a recursive update mechanism and method similarity to enhance the feedback on the execution results of both entities and methods. Experiments conducted on 186 real-world attack cases demonstrate that this model significantly improves the search efficiency compared with the current state-of-the-art baseline methods. Within the 0—2 000 step range, the average recall rate achieves a relative improvement of 18.24%. Notably, in the 0—250 step cold-start scenario, the average recall rate achieves a relative improvement of 86.28% compared to that achieved by the best baseline method. Furthermore, the ablation experiments indicate that each component of the proposed algorithm positively contributes to the overall performance, effectively reducing the cost of CTH.

Key words: Cyber Threat Hunting (CTH), proactive defense, Reinforcement Learning (RL), Large Language Model (LLM), Monte Carlo Tree Search (MCTS)

中图分类号: