[1] CHIB P S, SINGH P. Recent advancements in end-to-end autonomous driving using deep learning: a survey[J]. IEEE Transactions on Intelligent Vehicles, 2024, 9(1): 103-118. [2] 朱飞宇, 彭敦陆. YOLO-SP: 面向工业安全的局部多尺度目标检测检测网络[J]. 小型微型计算机系统, 2025, 46(6): 1435-1441. ZHU F Y, PENG D L. YOLO-SP: a local multi-scale object detection detection network for industrial safety[J]. Journal of Chinese Computer Systems, 2025, 46(6): 1435-1441. (in Chinese) [3] 俞益洲, 石德君, 马杰超, 等. 人工智能在医学影像分析中的应用进展[J]. 中国医学影像技术, 2019, 35(12): 1808-1812. YU Y Z, SHI D J, MA J C, et al. Advances in application of artificial intelligence in medical image analysis[J]. Chinese Journal of Medical Imaging Technology, 2019, 35(12): 1808-1812. (in Chinese) [4] CAI Z W, VASCONCELOS N. Cascade R-CNN: delving into high quality object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington D.C.,USA:IEEE Press,2018: 6154-6162. [5] REDMON J, DIVVALA S, GIRSHICK R, et al. You only look once: unified, real-time object detection[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C.,USA:IEEE Press,2016: 779-788. [6] LIU W, ANGUELOV D, ERHAN D, et al. SSD: single shot MultiBox detector[EB/OL].[2025-04-05]. https://arxiv.org/abs/1512.02325. [7] WANG X, ZHENG S F, YANG R, et al. Pedestrian attribute recognition: a survey[J]. Pattern Recognition, 2022, 121: 108220. [8] YE M, SHEN J B, LIN G J, et al. Deep learning for person re-identification: a survey and outlook[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(6): 2872-2893. [9] SIMONYAN K, ZISSERMAN A. Very deep convolutional networks for large-scale image recognition[EB/OL].[2025-04-05]. https://arxiv.org/abs/1409.1556. [10] SZEGEDY C, LIU W, JIA Y Q, et al. Going deeper with convolutions[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C.,USA:IEEE Press,2015: 1-9. [11] HE K M, ZHANG X Y, REN S Q, et al. Deep residual learning for image recognition[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C.,USA:IEEE Press,2016: 770-778. [12] GIRSHICK R, DONAHUE J, DARRELL T, et al. Rich feature hierarchies for accurate object detection and semantic segmentation[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Washington D.C.,USA:IEEE Press,2014: 580-587. [13] GIRSHICK R. Fast R-CNN[C]//Proceedings of the IEEE International Conference on Computer Vision (ICCV). Washington D.C.,USA:IEEE Press,2016: 1440-1448. [14] REN S Q, HE K M, GIRSHICK R, et al. Faster R-CNN: towards real-time object detection with region proposal networks[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6): 1137-1149. [15] CHENG X H, JIA M X, WANG Q, et al. A simple visual-textual baseline for pedestrian attribute recognition[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2022, 32(10): 6994-7004. [16] VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[EB/OL].[2025-04-05]. https://arxiv.org/abs/1706.03762. [17] DOSOVITSKIY A, BEYER L, KOLESNIKOV A, et al. An image is worth 16×16 words: transformers for image recognition at scale[EB/OL].[2025-04-05]. https://arxiv.org/abs/2010.11929. [18] DEVLIN J, CHANG M W, LEE K, et al. BERT: pre-training of deep bidirectional transformers for language understanding[EB/OL].[2025-04-05]. https://arxiv.org/abs/1810.04805. [19] 火久元, 苏泓瑞, 武泽宇, 等. 基于改进YOLOv8的道路交通小目标车辆检测算法[J]. 计算机工程, 2025, 51(1): 246-257. HUO J Y, SU H R, WU Z Y, et al. Road traffic small target vehicle detection algorithm based on improved YOLOv8[J]. Computer Engineering, 2025, 51(1): 246-257. (in Chinese) [20] LIU S, CHEN P F, WO AZĆG NIAK M. Image enhancement-based detection with small infrared targets[J]. Remote Sensing, 2022, 14(13): 3232. [21] 戚欣, 姜春雷. 人工智能助力智慧城市建设[J]. 智能建筑与智慧城市, 2017(9): 33-37. QI X, JIANG C L. Artificial intelligence boosts smart city construction[J]. Intelligent Building & City Information, 2017(9): 33-37. (in Chinese) [22] PAN X R, GE C J, LU R, et al. On the integration of self-attention and convolution[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C.,USA:IEEE Press,2022: 805-815. [23] WANG J W, XU C, YANG W, et al. A normalized Gaussian Wasserstein distance for tiny object detection[EB/OL].[2025-04-05]. https://arxiv.org/abs/2110.13389. [24] LI Y N, HUANG C, LOY C C, et al. Human attribute recognition by deep hierarchical contexts[EB/OL].[2025-04-05]. https://link.springer.com/content/pdf/10.1007/978-3-319-46466-4_41.pdf. [25] CAO Y R, HE Z J, WANG L J, et al. VisDrone-DET2021: the vision meets drone object detection challenge results[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW). Washington D.C.,USA:IEEE Press,2021: 2847-2854. [26] HE K M, GKIOXARI G, DOLLÁR P, et al. Mask R-CNN[C]//Proceedings of the IEEE International Conference on Computer Vision (ICCV). Washington D.C.,USA:IEEE Press,2017: 2980-2988. [27] CARION N, MASSA F, SYNNAEVE G, et al. End-to-end object detection with transformers[EB/OL].[2025-04-05]. https://arxiv.org/abs/2005.12872. [28] ZHU X Z, SU W J, LU L W, et al. Deformable DETR: deformable transformers for end-to-end object detection[EB/OL].[2025-04-05]. https://arxiv.org/abs/2010.04159. [29] LIU Z, LIN Y T, CAO Y, et al. Swin Transformer: hierarchical vision transformer using shifted windows[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Washington D.C.,USA:IEEE Press,2022: 9992-10002. [30] ZHAO Y A, LÜ W Y, XU S L, et al. DETRs beat YOLOs on real-time object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C.,USA:IEEE Press,2024: 16965-16974. [31] ZHANG N, PALURI M, RANZATO M, et al. PANDA: pose aligned networks for deep attribute modeling[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Washington D.C.,USA:IEEE Press,2014: 1637-1644. [32] 陈鸿昶, 吴彦丞, 李邵梅, 等. 基于行人属性分级识别的行人再识别[J]. 电子与信息学报, 2019, 41(9): 2239-2246. CHEN H C, WU Y C, LI S M, et al. Person re-identification based on attribute hierarchy recognition[J]. Journal of Electronics & Information Technology, 2019, 41(9): 2239-2246. (in Chinese) [33] ZHAO X, SANG L, DING G, et al. Grouping attribute recognition for pedestrian with joint recurrent learning[EB/OL].[2025-04-05]. https://www.ijcai.org/Proceedings/2018/0441.pdf. [34] 李娜, 武阳阳, 刘颖, 等. 基于多尺度注意力网络的行人属性识别算法[J]. 激光与光电子学进展, 2021, 58(4): 0410025. LI N, WU Y Y, LIU Y, et al. Pedestrian attribute recognition algorithm based on multi-scale attention network[J]. Laser & Optoelectronics Progress, 2021, 58(4): 0410025. (in Chinese) [35] DONG Q, GONG S G, ZHU X T. Multi-task curriculum transfer deep learning of clothing attributes[C]//Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV). Washington D.C.,USA:IEEE Press,2017: 520-529. [36] PHAM K, KAFLE K, LIN Z, et al. Learning to predict visual attributes in the wild[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C.,USA:IEEE Press,2021: 13013-13023. [37] CHEN K Y, JIANG X L, HU Y, et al. OvarNet: towards open-vocabulary object attribute recognition[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C.,USA:IEEE Press,2023: 23518-23527. [38] LI C L, YANG T, ZHU S J, et al. Density map guided object detection in aerial images[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). Washington D.C.,USA:IEEE Press,2020: 737-746. [39] LIU M J, WANG X H, ZHOU A J, et al. UAV-YOLO: small object detection on unmanned aerial vehicle perspective[J]. Sensors, 2020, 20(8): 2238. [40] LIU W J, QIANG J, LI X X, et al. UAV image small object detection based on composite backbone network[J]. Mobile Information Systems, 2022, 2022: 7319529. [41] XU C, DING J, WANG J W, et al. Dynamic coarse-to-fine learning for oriented tiny object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C.,USA:IEEE Press,2023: 7318-7328. [42] YANG C, HUANG Z H, WANG N Y. QueryDet: cascaded sparse query for accelerating high-resolution small object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C.,USA:IEEE Press,2022: 13658-13667. [43] LAI H Q, CHEN L Y, LIU W H, et al. STC-YOLO: small object detection network for traffic signs in complex environments[J]. Sensors, 2023, 23(11): 5307. [44] 张天鹏, 韩晶, 吕学强. 基于多任务学习的超分辨率辅助小目标检测[J]. 计算机工程, 2024, 50(9): 304-312. ZHANG T P, HAN J, LÜ X Q. Super-resolution-aided small-target detection based on multi-task learning[J]. Computer Engineering, 2024, 50(9): 304-312. (in Chinese) [45] 谢椿辉, 吴金明, 徐怀宇. 改进YOLOv5的无人机影像小目标检测算法[J]. 计算机工程与应用, 2023, 59(9): 198-206. XIE C H, WU J M, XU H Y. Small object detection algorithm based on improved YOLOv5 in UAV image[J]. Computer Engineering and Applications, 2023, 59(9): 198-206. (in Chinese) [46] TANG S Y, ZHANG S, FANG Y N. HIC-YOLOv5: improved YOLOv5 for small object detection[C]//Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). Washington D.C.,USA:IEEE Press,2024: 6614-6619. [47] WOO S, PARK J, LEE J Y, et al. CBAM: convolutional block attention module[EB/OL].[2025-04-05]. https://arxiv.org/abs/1807.06521. [48] WANG C Y, BOCHKOVSKIY A, LIAO H M. YOLOv7: trainable bag-of-freebies sets new state-of-the-art for real-time object detectors[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C.,USA:IEEE Press,2023: 7464-7475. [49] SONG G L, LIU Y, WANG X G. Revisiting the sibling head in object detector[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C.,USA:IEEE Press,2020: 11560-11569. [50] LIN T Y, MAIRE M, BELONGIE S, et al. Microsoft COCO: common objects in context[EB/OL].[2025-04-05]. https://arxiv.org/abs/1405.0312. [51] HU J, SHEN L, SUN G. Squeeze-and-excitation networks[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Washington D.C.,USA:IEEE Press,2018: 7132-7141. [52] RUBNER Y, TOMASI C, GUIBAS L J. The Earth mover’s distance as a metric for image retrieval[J]. International Journal of Computer Vision, 2000, 40(2): 99-121. [53] 石方炎. 人体检测与外观属性识别一体化算法研究[D]. 成都: 电子科技大学, 2020. SHI F Y. Research on integrated algorithm of human detection and appearance attribute recognition[D]. Chengdu: University of Electronic Science and Technology of China, 2020. (in Chinese) |