作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (10): 455-464. doi: 10.19678/j.issn.1000-3428.0252187

• 新一代网络与边缘计算 • 上一篇    

边缘计算中基于深度强化学习的任务安全卸载

张航1,2,3, 王劲松1,2,3   

  1. 1. 天津理工大学计算机科学与工程学院, 天津 300382;
    2. 智能计算机及软件新技术天津市重点实验室, 天津 300384;
    3. 计算机病毒防治技术国家工程实验室, 天津 300457
  • 收稿日期:2025-03-03 修回日期:2025-05-12 发布日期:2025-06-13
  • 作者简介:张航(CCF学生会员),男,博士研究生,主研方向为边缘计算、网络安全;王劲松(通信作者),教授、博士,E-mail:jswang@tjut.edu.cn。
  • 基金资助:
    天津市技术创新引导专项(22YDPYGX00040);天津市重点研发计划(23YFZCSN00240)。

Secure Task Offloading Based on Deep Reinforcement Learning in Edge Computing

ZHANG Hang1,2,3, WANG Jinsong1,2,3   

  1. 1. School of Computer Science and Engineering, Tianjin University of Technology, Tianjin 300382, China;
    2. Tianjin Key Laboratory of Intelligence Computing and Novel Software Technology, Tianjin 300384, China;
    3. National Engineering Laboratory for Computer Virus Prevention and Control Technology, Tianjin 300457, China
  • Received:2025-03-03 Revised:2025-05-12 Published:2025-06-13

摘要: 对于计算资源有限的用户设备(UD)而言,处理计算密集型任务较为困难。边缘计算通过将计算资源扩展到网络边缘来处理计算密集型任务,其关键功能之一便是计算任务的合理卸载。如何协调众多边缘节点的计算资源进行任务卸载,且在任务卸载过程中保障数据安全,是其重要挑战。因此,提出一种基于深度强化学习(DRL)的任务安全卸载方法。首先,构建边缘计算网络模型,并为其设计可变的安全防护机制,以适应性地保障数据安全;然后,将边缘计算网络模型和任务目标进行形式化,并将其转化为马尔可夫决策过程(MDP);最后,提出一种基于惩罚动作空间的DRL方法,给出最优的任务卸载策略。仿真结果表明,所提方案可以在进行安全防护的同时,降低时延和能源消耗成本,且始终保持零任务丢失率。

关键词: 边缘计算, 任务卸载, 马尔可夫决策过程, 安全机制, 深度强化学习

Abstract: The handling of computation-intensive tasks is challenging for User Devices (UD) with limited computing resources. Edge computing addresses this problem by extending computing resources to the network edge, and one of its key functions is rational computational offloading. However, coordinating the computing resources of numerous edge nodes for task offloading while ensuring data security throughout the process is a major challenge. To address this issue, this paper proposes a Deep Reinforcement Learning (DRL)-based secure task offloading method. First, an edge computing network model is constructed, and a variable security protection mechanism is designed to adaptively safeguard data. The edge computing network model and task objectives are formalized and transformed into a Markov Decision Process (MDP). Finally, a DRL method based on a penalized action space is proposed to derive the optimal task offloading strategy. Simulation results demonstrate that the proposed scheme can reduce latency and energy consumption costs while maintaining security protection, and it consistently achieves a zero task loss rate.

Key words: edge computing, task offloading, Markov Decision Process (MDP), security mechanism, Deep Reinforcement Learning (DRL)

中图分类号: