作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (10): 164-176. doi: 10.19678/j.issn.1000-3428.0252596

• 计算机视觉与图形图像处理 • 上一篇    

基于特征融合的多属性小目标检测

辛振伟1, 张粼1, 王宝龙2, 石少华2, 杨一鹏2, 程辉2, 陈涛1   

  1. 1. 复旦大学未来信息创新学院, 上海 200438;
    2. 西安昆仑工业(集团)有限责任公司, 陕西 西安 710043
  • 收稿日期:2025-06-09 修回日期:2025-08-17 发布日期:2026-09-29
  • 作者简介:辛振伟(CCF会员),男,硕士,主研方向为计算机视觉;张粼,博士;王宝龙、石少华,工程师;杨一鹏,助理工程师;程辉,工程师;陈涛(通信作者),教授,E-mail:eetchen@fudan.edu.cn。
  • 基金资助:
    国家重点研发计划(2022ZD0160101);上海市自然科学基金(23ZR1402900);上海市科学技术委员会"探索者计划"项目(24TS1401300);上海市科学技术重大项目(2021SHZDZX0103)。

Multi-Attribute Small Object Detection Based on Feature Fusion

XIN Zhenwei1, ZHANG Lin1, WANG Baolong2, SHI Shaohua2, YANG Yipeng2, CHENG Hui2, CHEN Tao1   

  1. 1. College of Future Information Technology, Fudan University, Shanghai 200438, China;
    2. Xi'an Kunlun Industrial (Group) Co., Ltd., Xi'an 710043, Shaanxi, China
  • Received:2025-06-09 Revised:2025-08-17 Published:2026-09-29

摘要: 目标检测是计算机视觉中的一项基础任务,旨在实现对图像中目标的定位和类别识别。多属性识别则是输入以目标为中心的图像,输出图像中目标的多个属性的分类结果。当前检测算法对小目标识别存在困难,而属性识别算法无法自行定位目标。为此,提出一种多属性小目标检测框架MA-YOLO,以同时完成小目标检测和多属性识别两个任务。首先,该框架通过属性解耦模块(ADM)进行特征解耦,并添加多个并行的属性分类头;其次,为了实现属性间的特征交互,设计一种简单但有效的属性融合策略,通过桥接特征生成模块(BFGM)生成桥接特征,实现骨干网络提取到的特征和后续解耦出来的所有属性特征的过渡,利用属性融合模块(AFM)得到融合后的各个属性特征。针对小目标检测问题,采用一种卷积和自注意力混合(ACmix)模块,以增强网络的特征提取能力。此外,将边界框建模为二维高斯分布,使用归一化Wasserstein距离(NWD)度量。实验结果表明,在多属性数据集WIDER Attribute的测试集和验证集以及小目标数据集VisDrone2019的验证集上,MA-YOLO的mAP@0.5分别达到84.1%、85.3%和48.2%,大量消融和可视化实验也验证了MA-YOLO的有效性。

关键词: 深度学习, 计算机视觉, 目标检测, 多属性, 小目标

Abstract: Object detection is a fundamental task in computer vision, which aims to locate and classify objects within an image. By contrast, multi-attribute recognition involves inputting an image centered around an object and outputting the classification results of multiple attributes of the object within the image. Current detection algorithms face challenges in recognizing small objects, while attribute recognition algorithms fail to locate objects. To address these issues, this paper proposes a multi-attribute small object detection framework, MA-YOLO, that performs small object detection and multi-attribute recognition simultaneously. First, an Attribute Decoupling Module (ADM) decouples attributes of the object and adds multiple parallel attribute classification heads. Second, a simple yet effective attribute-fusion strategy is used to facilitate feature interactions among the attributes. This strategy generates bridging features through a Bridge Feature Generation Module (BFGM) to bridge the gap between the features extracted by the backbone network and all decoupled attribute features, and it utilizes an Attribute Fusion Module (AFM) to obtain fused attribute features. For small object detection, a convolution and self-attention hybrid module, ACmix, is adopted to enhance the feature extraction capability of the network. Additionally, bounding boxes are modeled as two-dimensional Gaussian distributions and measured using the Normalized Wasserstein Distance (NWD). Experimental results show that MA-YOLO achieves mAP@0.5 values of 84.1% and 85.3% on the test and validation sets of the multi-attribute dataset WIDER Attribute and 48.2% as the validation set of the small object dataset VisDrone2019. Extensive ablation and visualization experiments validate the effectiveness of MA-YOLO.

Key words: deep learning, computer vision, object detection, multi-attribute, small object

中图分类号: