作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (10): 73-80. doi: 10.19678/j.issn.1000-3428.0252115

• 计算智能与模式识别 • 上一篇    

基于双向构造策略的神经组合优化模型

王朝扬, 孙未未   

  1. 复旦大学计算机科学技术学院, 上海 200433
  • 收稿日期:2025-02-11 修回日期:2025-04-11 发布日期:2025-05-20
  • 作者简介:王朝扬,男,硕士研究生,主研方向为神经组合优化;孙未未(通信作者),教授,E-mail:wwsun@fudan.edu.cn。
  • 基金资助:
    国家自然科学基金(62172107)。

Neural Combinatorial Optimization Model Based on Bidirectional Construction Strategy

WANG Chaoyang, SUN Weiwei   

  1. School of Computer Science, Fudan University, Shanghai 200433, China
  • Received:2025-02-11 Revised:2025-04-11 Published:2025-05-20

摘要: 组合优化问题在物流路径规划等领域具有重要应用价值,但其解空间随问题规模呈指数级扩张,导致传统方法面临严峻挑战。近年来基于强化学习的神经组合优化方法能够在保持较短求解耗时的同时,使解质量接近传统求解器的水平。主流方法POMO(Policy Optimization with Multiple Optima)通过对称性优化增强了训练稳定性,但其单向序列生成机制仍存在双重局限:一方面,传统构造式方法难以充分挖掘问题对称性特征;另一方面,终点信息无法有效参与远端节点的决策过程。针对这一问题,提出了基于双向构造策略(BCS)的BCS-POMO模型,通过起点与终点双向并行构造解序列,动态选择更有把握的扩展方向,避免模型因单向构造而陷入两难的抉择之中。该模型利用构造序列对称性实现权重参数共享,并通过批量并行计算提升效率。实验表明,BCS-POMO有效强化了终点信息在构造过程中的决策辅助作用,在旅行商问题(TSP)和有容量约束的车辆路径问题(CVRP)上分别使误差降低了16%和18%,验证了双向构造策略对终点信息利用的有效性和对称性建模的优势。

关键词: 组合优化问题, 神经组合优化方法, 注意力机制, 强化学习, 旅行商问题, 有容量约束的车辆路径问题

Abstract: Combinatorial optimization problems have important applications in areas such as logistics path planning; however, their solution space exponentially expands with the problem size, leading to severe challenges for traditional methods. In recent years, neural combinatorial optimization methods based on reinforcement learning have been able to achieve solution quality close to that of traditional solvers while keeping the solution computation time short. The mainstream method, called Policy Optimization with Multiple Optima (POMO), enhances the training stability through symmetry optimization; however, its unidirectional sequence generation mechanism still suffers from two limitations: the traditional constructive method cannot fully exploit the symmetry features of the problem easily and the endpoint information cannot effectively participate in the decision-making process of the remote node. To address this problem, this paper proposes a Bidirectional Construction Strategy (BCS)-based POMO model, named BCS-POMO, which dynamically selects the extension direction with higher confidence by constructing the solution sequence in parallel from the start and end points, avoiding models that are caught in a dilemma owing to unidirectional constructions. The model exploits the symmetry of the construction sequence to achieve weight parameter sharing and improves efficiency through batch parallel computation. Experimental results show that the BCS-POMO effectively reinforces the role of endpoint information as a decision aid in the construction process, which reduces the error by 16% and 18% for the Traveling Salesman Problem (TSP) and the Capacitated Vehicle Routing Problem (CVRP), respectively, verifying the effectiveness of the bidirectional construction strategy in exploiting endpoint information and the advantages of symmetry modelling.

Key words: combinatorial optimization problem, neural combinatorial optimization method, attention mechanism, reinforcement learning, Traveling Salesman Problem (TSP), Capacitated Vehicle Routing Problem (CVRP)

中图分类号: