作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (10): 116-128. doi: 10.19678/j.issn.1000-3428.0070761

• 计算智能与模式识别 • 上一篇    

基于最优传输距离正则化和近邻聚类方法的开放集域适应

田青1,2,3, 申珺妤1, 郁江森1   

  1. 1. 南京信息工程大学软件学院, 江苏 南京 210044;
    2. 南京信息工程大学无锡研究院, 江苏 无锡 214000;
    3. 南京航空航天大学模式分析与机器智能工业和信息化部重点实验室, 江苏 南京 211106
  • 收稿日期:2024-12-30 修回日期:2025-03-31 发布日期:2025-05-22
  • 作者简介:田青(CCF高级会员),男,教授、博士、博士生导师,主研方向为机器学习、模式识别、计算机视觉,E-mail:tianqing@nuist.edu.cn;申珺妤、郁江森,硕士研究生。
  • 基金资助:
    国家自然科学基金(62176128);江苏省自然科学基金(BK20231143);中央高校基本科研业务费专项资金(NJ2023032);江苏省"青蓝工程"人才计划。

Open-Set Domain Adaptation Based on Optimal Transport Distance Regularization and Neighbor Clustering Method

TIAN Qing1,2,3, SHEN Junyu1, YU Jiangsen1   

  1. 1. School of Software, Nanjing University of Information Science and Technology, Nanjing 210044, Jiangsu, China;
    2. Wuxi Institute of Technology, Nanjing University of Information Science and Technology, Wuxi 214000, Jiangsu, China;
    3. MIIT Key Laboratory of Pattern Analysis and Machine Intelligence, Nanjing University of Aeronautics and Astronautics, Nanjing 211106, Jiangsu, China
  • Received:2024-12-30 Revised:2025-03-31 Published:2025-05-22

摘要: 无监督域适应(UDA)旨在将知识从标记的源域迁移到未标记的目标域,从而提高目标域模型的性能。然而,传统的UDA方法假设源域和目标域的类别空间完全一致,导致无法处理目标域中存在的未知类别,这限制了其在实际场景中的应用。开放集域适应(OSDA)通过引入对未知类别的识别来解决这一问题,但如何有效减少域间差异和类别不平衡对模型性能的负面影响,仍是一个重要挑战。现有的OSDA方法往往忽略了域特定特征,并简单地对域差异直接进行最小化,这可能导致类别之间的边界不清晰并削弱模型的泛化能力。为了解决这一问题,提出基于最优传输距离正则化(OTR)和近邻聚类方法的开放集域适应(OTRNC)方法。该方法通过OTR来最大化高、低置信度样本组之间的分布距离,减少未知类别对域适应过程的干扰。然后利用动态近邻检索和不变特征学习,减少目标域的类内变化,增强特征的泛化能力。实验结果表明,OTRNC在多个基准数据集上均表现出色,具有一定的有效性。

关键词: 开放集域适应, 最优传输距离正则化, 动态近邻检索, 不变特征学习, 近邻聚类

Abstract: Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain, thereby improving the performance of the target domain model. However, traditional UDA methods assume that the category spaces of the source and target domains are identical; hence, they cannot handle unknown categories present in the target domain, which limits their applicability in real-world scenarios. Open-Set Domain Adaptation (OSDA) addresses this issue by introducing the recognition of unknown categories. Nevertheless, effectively mitigating the negative impact of inter-domain discrepancies and class imbalances on the model performance remains a significant challenge. Existing OSDA methods tend to overlook domain-specific features and directly minimize domain discrepancies, which may lead to ambiguous decision boundaries between classes and weaken the generalization ability of the model. To address this problem, this paper proposes an open-set domain adaptation method based on Optimal Transport Distance Regularization (OTR) and neighborhood clustering, called OTRNC. Specifically, the method employs OTR to maximize the distributional distance between high-confidence and low-confidence sample groups, thereby reducing the interference of unknown categories in the domain adaptation process. Subsequently, it leverages dynamic neighbor retrieval and invariant feature learning to reduce intra-class variations in the target domain and enhance the generalization capability of the learned features. Experimental results demonstrate that OTRNC achieves superior performance across multiple benchmark datasets, thereby validating its effectiveness.

Key words: Open-Set Domain Adaptation (OSDA), Optimal Transport Distance Regularization (OTR), dynamic neighbor retrieval, invariant feature learning, neighbor clustering

中图分类号: