作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (10): 242-251. doi: 10.19678/j.issn.1000-3428.0252150

• 计算机视觉与图形图像处理 • 上一篇    

基于提示信息传输框架的图文行人重识别

耿霞1, 林贤文1, 杨治2   

  1. 1. 江苏大学计算机科学与通信工程学院, 江苏 镇江 212013;
    2. 江苏大学管理学院, 江苏 镇江 212013
  • 收稿日期:2025-02-20 修回日期:2025-04-30 发布日期:2025-06-03
  • 作者简介:耿霞,女,副教授、博士,主研方向为模式识别、人工智能、生物信息学、机器视觉,E-mail:gengxia@ujs.edu.cn;林贤文,硕士研究生;杨治,博士研究生。
  • 基金资助:
    国家自然科学基金面上项目(62276116)。

Prompt-Based Information Transfer Framework for Image-Text Person Re-Identification

GENG Xia1, LIN Xianwen1, YANG Zhi2   

  1. 1. School of Computer Science and Communication Engineering, Jiangsu University, Zhenjiang 212013, Jiangsu, China;
    2. School of Management, Jiangsu University, Zhenjiang 212013, Jiangsu, China
  • Received:2025-02-20 Revised:2025-04-30 Published:2025-06-03

摘要: 在基于图文的行人重识别任务中,基于图文预训练模型的参数初始化已成为主流范式,这有效突破了单模态模型因跨模态信息缺失导致的特征对齐瓶颈。现有方法聚焦于挖掘图像-文本联合嵌入空间中不同尺度下的语义特征进行优化,但新对齐范式的引入易使原模型在微调过程中陷入局部最优。为了解决上述问题,提出一种基于提示的信息传输(PIT)框架,通过在单模态编码器和跨模态图像-文本编码器的原始前向过程中嵌入跨模态提示标识符,促进早期特征融合,隐式地引导模型更加聚焦于模态不变的信息。PIT框架包含基于跨模态提示的对比学习(CL)损失和提示训练策略(PTS)。基于跨模态提示的CL损失旨在通过约束图文特征之间的相似度,构建兼具模态内区分度与模态间语义一致性的共享特征嵌入空间。PTS可以视为一种自蒸馏方法,通过将无提示特征与基准真相产生的伪目标视为另一种行人图文对的特征视图,监督跨模态提示特征的训练过程,使最终学习到的特征嵌入相较于无提示特征包含更丰富的多模态信息。实验结果表明,PIT在完全微调的基础上仅需要添加0.61×106的参数,在CUHK-PEDES、ICFG-PEDES和RSTPReid数据集上所提模型相比基线模型的Rank-1提高了1.48、1.5 和1.55百分点。

关键词: 跨模态, 图文的行人重识别, 提示学习, BLIP模型, 自蒸馏学习

Abstract: In Image—Text person Re-IDentification (ITReID) tasks, initializing models with parameters from pretraining models has become a mainstream paradigm that effectively alleviates the feature alignment bottleneck of single-modal models caused by a lack of cross-modal information. Existing methods focus on mining semantic features at different scales in an image-text joint embedding space for optimization. However, the introduction of the new alignment paradigm is prone to causing the pretraining model to fall to a local minimum during fine-tuning. To address these issues, this paper proposes a Prompt-based Information Transfer (PIT) framework. By introducing cross-modal prompt tokens into the original forward process of the single-modal encoder and cross-modal image-text encoder, the proposed framework promotes early feature fusion and implicitly guides the pretraining model to focus more on modal-invariant information. PIT includes a prompt-based Contrastive Learning (CL) loss and Prompt Training Strategy (PTS). Prompt-based contrastive loss aims to construct a shared feature-embedding space with both intra-modal discrimination and inter-modal semantic consistency by constraining the similarity between graphic and text features. The PTS can be regarded as a form of self-distillation that treats the pseudo-targets generated by nonprompt features and ground truth as another view of image-text pairs, supervises the training process, and makes the learned embeddings contain richer multimodal information. Only 0.61×106 additional parameters are introduced based on fine-tuning, and PIT achieves Rank-1 improvements of 1.48, 1.5, and 1.55 percentage points on three public datasets.

Key words: cross-modality, Image-Text person Re-IDentification (ITReID), prompt learning, BLIP model, self-distillation learining

中图分类号: