作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程

• •    

基于深度强化学习的分层跨域资源分配方法

  • 发布日期:2026-07-22

A Hierarchical Cross-Domain Resource Allocation Method Based on Deep Reinforcement Learning

  • Published:2026-07-22

摘要: 高效的跨域资源分配是空天地一体化网络实现资源优化配置与服务质量保障的基础。然而,受限于空天地一体化网络固有的多层高度复杂拓扑和大量差异化网络切片服务请求,传统扁平化架构和静态数学优化模型在面对动态复杂问题时,难以实现跨域资源的高效与动态分配。因此提出一种基于深度强化学习的分层跨域资源分配方法(DLCDS),旨在最大化系统效用的同时兼顾底层物理资源的高效利用。首先,设计了分层协作的跨域框架,建立全局域层与本地域层的双层协作机制。其次,在全局域层面,提出了基于双延迟深度确定性策略梯度的服务等级协议(SLA)分解算法,该算法借助连续动作空间下的确定性策略实现了SLA约束在各子域间的规划与分解;在本地域层面,设计了基于近端策略优化的资源分配算法,该算法构建SLA映射模型实现了从服务目标到资源需求的转换,采用随机策略的探索优势提升域内资源分配效率。最后,实验结果表明,相比于其他基准算法,所提算法资源利用率平均提升12%,系统效用平均提升19%。

Abstract: Efficient cross-domain resource allocation is the foundation for achieving optimized resource distribution and ensuring service quality in the Space-Air-Ground Integrated Network (SAGIN). However, constrained by the inherent multi-layer highly complex topology of SAGIN and a massive number of differentiated network slicing service requests, traditional flat architectures and static mathematical optimization models struggle to achieve efficient and dynamic cross-domain resource allocation in the face of dynamic and complex problems. Therefore, this paper proposes a hierarchical cross-domain resource allocation method based on deep reinforcement learning (DLCDS), aiming to maximize system utility while ensuring the efficient utilization of underlying physical resources. Firstly, this paper designs a hierarchical collaborative cross-domain framework, establishing a dual-layer collaborative mechanism between the global domain layer and the local domain layer. Secondly, at the global domain layer, a service level agreement (SLA) decomposition algorithm based on the Twin Delayed Deep Deterministic Policy Gradient is proposed. This algorithm leverages deterministic policies within continuous action spaces to achieve the planning and decomposition of SLA constraints across subdomains. At the local domain level, a resource allocation algorithm based on Proximal Policy Optimization is designed. This algorithm constructs an SLA mapping model to achieve the conversion from service objectives to resource requirements, and employs the exploratory advantage of stochastic policies to enhance resource allocation efficiency within the domain. Finally, experimental results demonstrate that, compared to other benchmark algorithms, this method achieves an average 12% improvement in resource utilization and a 19% increase in system utility.