Author Login Chief Editor Login Reviewer Login Editor Login Remote Office

Computer Engineering ›› 2026, Vol. 52 ›› Issue (7): 421-433. doi: 10.19678/j.issn.1000-3428.0070379

• Interdisciplinary Integration and Engineering Applications • Previous Articles     Next Articles

Joint Scheduling Optimization of Distributed Resources in a Container Terminal Yard Based on Actor—Critic Deep Reinforcement Learning

DONG Liangcai*(), YUE Hanci, WANG Weijuan   

  1. Logistics Engineering College, Shanghai Maritime University, Shanghai 201306, China
  • Received:2024-09-14 Revised:2024-11-01 Online:2026-07-15 Published:2026-07-04
  • Contact: DONG Liangcai

基于Actor-Critic深度强化学习的堆场分布式资源联合调度优化

董良才*(), 岳涵词, 王维娟   

  1. 上海海事大学物流工程学院, 上海 201306
  • 通讯作者: 董良才
  • 作者简介:

    董良才, 男, 副教授、博士, 主研方向为物流系统智能决策

    岳涵词, 硕士研究生

    王维娟, 硕士研究生

  • 基金资助:
    国家自然科学基金(52472435); 上海市科学技术委员会项目(22ZR1427700); 上海市教育科学项目(B2023003)

Abstract:

As export container throughput continues to increase in container terminal yards, ensuring the efficient management and scheduling of terminal resources has become crucial for enhancing terminal competitiveness. Considering an automated container terminal yard as the research object, this study proposes an optimization method based on Deep Reinforcement Learning (DRL) to address the issues of storage slot allocation and yard crane scheduling for export containers. A comprehensive analysis of the container terminal yard operation system is performed, and a multi-objective optimization model is constructed. This model primarily aims to reduce the yard crane operation time, internal truck waiting time, and number of container relocations within the yard blocks, while also considering constraints such as the safe distance between yard cranes and balanced workload distribution. To reduce the complexity of solving the model, an Actor—Critic algorithm based on DRL is proposed. Numerical examples of different scales are designed. Through a comparative analysis with the results of the Genetic Algorithm (GA) and exact solutions from CPLEX, the advantages of the Actor—Critic algorithm in terms of solution speed and quality are demonstrated. Experiments reveal that the proposed algorithm can quickly and accurately solve small-scale problems and obtain near-optimal solutions for large-scale problems, exhibiting significant superiority over the GA in solving large-scale problems. Comparative experiments are conducted to further investigate the impact of the number of yard cranes and zoning-balanced mixed stacking strategy on the optimization results. The analysis results indicate that the zoning-balanced mixed stacking strategy outperforms traditional stacking strategies in terms of balancing yard crane workloads and reducing idle time.

Key words: storage space allocation, yard crane scheduling, Deep Reinforcement Learning (DRL), Actor—Critic, multi-objective optimization

摘要:

面对日益增长的集装箱吞吐量, 如何高效地管理和调度码头资源, 成为提升码头竞争力的关键所在。以自动化集装箱码头堆场为研究对象, 针对出口箱的箱位分配和场桥调度问题, 提出一种基于深度强化学习(DRL)的优化方法。对集装箱码头堆场作业系统进行全面分析, 构建一个多目标优化模型, 该模型以减少场桥作业时间、集卡等待时间和箱区翻箱量为主要目标, 同时兼顾场桥间的安全距离和作业量均衡等约束条件。为了降低模型求解的复杂性, 提出基于DRL的Actor-Critic算法。设计不同规模的算例, 通过与遗传算法(GA)的结果和CPLEX的精确解之间的对比分析, 展示Actor-Critic算法在求解速度和解的质量上的优势。实验结果表明, 该算法在小规模问题上能够快速准确地求解, 在大规模问题上也能获得近似最优解, 且在大规模问题的求解上, 相较于GA展现了显著的优越性。通过对比实验, 进一步探讨了场桥数量和分区平衡混堆策略对优化结果的影响。分析结果表明, 分区平衡混堆策略在均衡场桥作业量、减少场桥空闲时间方面优于传统堆存策略。

关键词: 箱位分配, 场桥调度, 深度强化学习, Actor-Critic, 多目标优化