[1] LI S, XIAO T, LI H S, et al. Person search with natural language description[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Washington D.C., USA: IEEE Press, 2017: 1970-1979. [2] HAN X Y, ZHONG X, HUANG W X, et al. See what you seek: semantic contextual integration for cloth-changing person re-identification[EB/OL].[2025-01-15]. https://arxiv.org/pdf/2412.01345. [3] RADFORD A, KIM J W, HALLACY C, et al. Learning transferable visual models from natural language supervision[C]//Proceedings of International Conference on Machine Learning. New York, USA: [s.n], 2021: 8748-8763. [4] LI J N, SELVARAJU R R, GOTMARE A D, et al. Align before fuse: vision and language representation learning with momentum distillation[EB/OL].[2025-01-15]. https://arxiv.org/pdf/2107.07651. [5] LI J N, LI D X, XIONG C M, et al. BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation[EB/OL]. [2025-01-15]. https://arxiv.org/pdf/2201.12086. [6] JIANG D, YE M. Cross-modal implicit relation reasoning and aligning for text-to-image person retrieval[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C., USA: IEEE Press, 2023: 2787-2797. [7] YAN S L, DONG N, ZHANG L Y, et al. CLIP-driven fine-grained text-image person re-identification[J]. IEEE Transactions on Image Processing, 2023, 32: 6032-6046. [8] LIU Y T, LI Y W, LIU Z M, et al. CLIP-based synergistic knowledge transfer for text-based person retrieval[C]//Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Washington D.C., USA: IEEE Press, 2024: 7935-7939. [9] SUN J T, FEI H, ZHENG Z D, et al. From data deluge to data curation: a filtering-WoRA paradigm for efficient text-based person search[EB/OL]. [2025-01-15]. https://arxiv.org/pdf/2404.10292. [10] YANG S Y, ZHOU Y N, ZHENG Z D, et al. Towards unified text-based person retrieval: a large-scale multi-attribute and language search benchmark[C]//Proceedings of the 31st ACM International Conference on Multimedia. New York, USA: ACM Press, 2023: 4492-4501. [11] HE K M, FAN H Q, WU Y X, et al. Momentum contrast for unsupervised visual representation learning[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C., USA: IEEE Press, 2020: 9726-9735. [12] 孟琪翔, 高志霖,王劲滔,等. 面向公共场所敏感目标与人体异常行为协同识别网络[J]. 光电工程, 2025, 52(8): 250131. MENG Q X, GAO Z L, WANG J T, et al. Collaborative detection network for sensitive targets and abnormal human behaviour in public places[J]. Opto-Electronic Engineering, 2025, 52(8): 250131. (in Chinese) [13] DING Z F, DING C X, SHAO Z Y, et al. Semantically self-aligned network for text-to-image part-aware person re-identification[EB/OL]. [2025-01-15]. https://arxiv.org/pdf/2107.12666. [14] ZHU A C, WANG Z J, LI Y F, et al. DSSL: deep surroundings-person separation learning for text-based person retrieval[C]//Proceedings of the 29th ACM International Conference on Multimedia. New York, USA: ACM Press, 2021: 209-217. [15] ZHENG Z D, ZHENG L, GARRETT M, et al. Dual-path convolutional image-text embeddings with instance loss[J]. ACM Transactions on Multimedia Computing, Communications, and Applications, 2020, 16(2): 1-23. [16] 姜定, 叶茫. 面向跨模态文本到图像行人重识别的Transformer网络[J]. 中国图象图形学报, 2023, 28(5): 1384-1395. JIANG D, YE M. Transformer network for cross-modal text-to-image person re-identification[J]. Journal of Image and Graphics, 2023, 28(5): 1384-1395. (in Chinese) [17] YAN S L, DONG N, LIU J, et al. Learning comprehensive representations with richer self for text-to-image person re-identification[C]//Proceedings of the 31st ACM International Conference on Multimedia. New York, USA: ACM Press, 2023: 6202-6211. [18] WANG Z J, ZHU A C, XUE J Y, et al. CAIBC: capturing all-round information beyond color for text-based person retrieval[C]//Proceedings of the 30th ACM International Conference on Multimedia. New York, USA: ACM Press, 2022: 5314-5322. [19] 王晋溪, 鲁鸣鸣. 基于场景图知识的文本到图像行人重识别[J]. 模式识别与人工智能, 2024, 37(11): 947-959. WANG J X, LU M M. Scene graph knowledge based text-to-image person re-identification[J]. Pattern Recognition and Artificial Intelligence, 2024, 37(11): 947-959. (in Chinese) [20] BAI Y, CAO M, GAO D M, et al. RaSa: relation and sensitivity aware representation learning for text-based person search[C]//Proceedings of the 32nd International Joint Conference on Artificial Intelligence. New York, USA: ACM Press, 2023: 555-563. [21] KHATTAK M U, RASHEED H, MAAZ M, et al. MaPLe: multi-modal prompt learning[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C., USA: IEEE Press, 2023: 19113-19122. [22] LIN D X, PENG Y X, MENG J K, et al. Cross-modal adaptive dual association for text-to-image person retrieval[J]. IEEE Transactions on Multimedia, 2024, 26: 6609-6620. [23] YE M, SHEN J B, LIN G J, et al. Deep learning for person re-identification: a survey and outlook[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 44(6): 2872-2893. [24] CHEN Y H, ZHANG G Q, LU Y J, et al. TIPCB: a simple but effective part-based convolutional baseline for text-based person search[J]. Neurocomputing, 2022, 494: 171-181. [25] FUJII T, TARASHIMA S. BiLMa: bidirectional local-matching for text-based person re-identification[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW). Washington D.C., USA: IEEE Press, 2023: 2778-2782. [26] QIN Y, CHEN Y K, PENG D Z, et al. Noisy-correspondence learning for text-to-image person re-identification[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Washington D.C., USA: IEEE Press, 2024: 27187-27196. [27] CAO M, BAI Y, ZENG Z Y, et al. An empirical study of CLIP for text-based person search[C]// Proceedings of the AAAI Conference on Artificial Intelligence. [S.l.]: AAAI Press, 2024: 465-473. [28] ERGASTI A, FONTANINI T, FERRARI C, et al. MARS: paying more attention to visual attributes for text-based person search[J]. ACM Transactions on Multimedia Computing Communications and Applications, 2025, 21(10):1-22. [29] DENG Y C, HU Z P, HAN J K, et al. DualFocus: integrating plausible descriptions in text-based person re-identification[EB/OL]. [2025-01-15]. https://arxiv.org/pdf/2405.07459. |