作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (10): 263-274. doi: 10.19678/j.issn.1000-3428.0252029

• 计算机视觉与图形图像处理 • 上一篇    

基于改进YOLOv8与ByteTrack的航拍多目标跟踪算法

郑明宇1, 邵慧超2, 邵延华1, 楚红雨1   

  1. 1. 西南科技大学信息工程学院, 四川 绵阳 621010;
    2. 立得空间信息技术股份有限公司, 湖北 武汉 430070
  • 收稿日期:2025-01-09 修回日期:2025-03-11 发布日期:2025-05-07
  • 作者简介:郑明宇(CCF会员),女,硕士研究生,主研方向为目标检测、目标跟踪;邵慧超,高级工程师;邵延华(CCF会员),副教授、博士;楚红雨(通信作者),教授、博士,E-mail:49456338@qq.com。
  • 基金资助:
    国家自然科学基金(61601382,62106209)。

Multi-Target Tracking Algorithm for Aerial Photography Based on Improved YOLOv8 and ByteTrack

ZHENG Mingyu1, SHAO Huichao2, SHAO Yanhua1, CHU Hongyu1   

  1. 1. School of Information Engineering, Southwest University of Science and Technology, Mianyang 621010, Sichuan, China;
    2. Leador Spatial Information Technology Co., Ltd., Wuhan 430070, Hubei, China
  • Received:2025-01-09 Revised:2025-03-11 Published:2025-05-07

摘要: 目标检测与多目标跟踪技术日益成熟,但在复杂场景下执行航拍多目标跟踪任务时,目标尺寸小、尺寸变化大、遮挡等问题仍会导致检测与跟踪效果不理想。为此,提出一种基于改进YOLOv8与ByteTrack的航拍多目标跟踪算法(YBTrack)。构建MSA-YOLO(Multi-Scale Attention YOLO)检测器,设计空间-深度卷积替换YOLOv8原有卷积,将空间信息转换为通道维度,有效保留目标细节,减少多尺度特征图融合过程中信息丢失导致的漏检、误检。同时设计轻量加速空间-通道注意力模块,用于颈部卷积,降低计算复杂度,并作为检测头前的特征细化模块,进一步增强算法对目标特征信息提取的能力。为提高跟踪效果,对ByteTrack跟踪模型进行优化,设计空间-外观相似度矩阵(ASM),提升模型区分相似目标的性能,并提出目标校正函数,减少卡尔曼滤波器产生的误差积累,降低目标偏移、丢失率。将MSA-YOLO与优化后的ByteTrack结合,开展多目标跟踪实验。其中,MSA-YOLO在VisDrone2019-DET数据集上平均精度均值(mAP@0.5)提高了9.4百分点;多目标跟踪算法在VisDrone2019-MOT、MOT17数据集上多目标跟踪准确率(MOTA)分别提升了11.2和8.3百分点,ID调和均值(IDF1)分别提升了8.9和7.4百分点,实验结果表明本文所提方法跟踪效果显著。此外,与其他多目标跟踪算法的对比实验也证明了本文算法的优越性。

关键词: YOLOv8, ByteTrack, 多目标跟踪, 特征信息, 卡尔曼滤波器

Abstract: Object detection and multi-target tracking technologies are becoming increasingly mature. However, when performing aerial multi-target tracking tasks in complex scenarios, issues such as small target size, large size variation, and occlusion lead to unsatisfactory detection and tracking performances. Therefore, this paper proposes an aerial multi-target tracking algorithm based on an improved YOLOv8 and ByteTrack method called YBTrack. First, a detector called Multi-Scale Attention YOLO (MSA-YOLO) is constructed. The original convolution in YOLOv8 is replaced with a space-depth convolution that transforms spatial information into channel dimensions, effectively preserves target details, and reduces missed and false detections caused by information loss during multiscale feature map fusion. Simultaneously, a lightweight accelerated space-channel attention module is designed for neck convolution to reduce computational complexity. This module also acts as a feature refinement module before the detection head, further enhancing its ability to extract target feature information. Next, the ByteTrack tracking model is optimized to improve the tracking performance. A Spatial-Appearance Similarity Matrix (ASM) is designed to enhance the model's ability to distinguish similar targets. In addition, a target correction function is proposed to reduce the error accumulation of the Kalman filter and decrease the target offset and loss rates. Finally, MSA-YOLO and the optimized ByteTrack are combined for multi-target tracking experiments. MSA-YOLO achieves a 9.4 percentage points improvement in mAP@0.5 on the VisDrone2019-DET dataset. The multi-target tracking algorithm improves the Multiple Object Tracking Accuracy (MOTA) by 11.2 and 8.3 percentage points and the Identification F1 Score (IDF1) by 8.9 and 7.4 percentage points on the VisDrone2019-MOT and MOT17 datasets, respectively. Experimental results demonstrate the excellent tracking performance of the proposed method. Furthermore, comparison experiments with other multitarget tracking algorithms confirm the superiority of the proposed algorithm.

Key words: YOLOv8, ByteTrack, multiple target tracking, feature information, Kalman filter

中图分类号: