Author Login Chief Editor Login Reviewer Login Editor Login Remote Office

Computer Engineering ›› 2026, Vol. 52 ›› Issue (9): 204-217. doi: 10.19678/j.issn.1000-3428.0070406

• Computer Vision and Image Processing • Previous Articles     Next Articles

Cross-View Geo-Localization Based on Dynamic Metric Learning in UAV Scenarios

WANG Xinxin1, HU Haifeng1,*(), ZHANG Suofei2, ZHOU Feifei3, GONG Rui3   

  1. 1. School of Communications and Information Engineering, Nanjing University of Posts and Telecommunications, Nanjing 210003, Jiangsu, China
    2. School of Internet of Things, Nanjing University of Posts and Telecommunications, Nanjing 210003, Jiangsu, China
    3. China Telecom Corporation Limited Jiangsu Branch, Nanjing 210037, Jiangsu, China
  • Received:2024-09-24 Revised:2025-02-12 Online:2026-09-15 Published:2026-09-01
  • Contact: HU Haifeng

无人机场景下基于动态度量学习的跨视角地理定位

王鑫鑫1, 胡海峰1,*(), 张索非2, 周飞飞3, 龚锐3   

  1. 1. 南京邮电大学通信与信息工程学院, 江苏 南京 210003
    2. 南京邮电大学物联网学院, 江苏 南京 210003
    3. 中国电信股份有限公司江苏分公司, 江苏 南京 210037
  • 通讯作者: 胡海峰
  • 作者简介:

    王鑫鑫, 男, 硕士研究生, 主研方向为计算机视觉、地理定位

    胡海峰(通信作者), 教授、博士

    张索非, 博士

    周飞飞, 硕士研究生

    龚锐, 硕士研究生

  • 基金资助:
    国家自然科学基金(62371245)

Abstract:

Previous studies on cross-view geolocalization have primarily focused on determining whether a query image accurately corresponds to a specific geographic location within a predefined gallery. However, this research paradigm often overlooks the extensive multiscale structural information inherent in the geographic space. To achieve more robust localization, a model must not only capture local architectural details but also understand the spatial relationships among targets reflected through building clusters and environmental features, thereby improving the localization accuracy across different spatial scales. To address these challenges, a multiscale cross-view geo-localization task is proposed and a Multi-Level Campus (ML-Campus) dataset is constructed specifically for this task. The ML-Campus dataset comprises multiview, multisource building images, each annotated with multiscale labels to reflect correlations and continuity across different spatial scales. Based on this dataset, an empirical evaluation of existing cross-view geo-localization methods is conducted, which is used as a benchmark to measure their performance in this context. To further enhance model performance, the proposed Cross-View HAPPIER (CV-HAPPIER) method is employed for training, which strengthens the model's feature representation capabilities across different spatial scales. Extensive experimental results on the ML-Campus dataset demonstrate that the CV-HAPPIER method significantly improves the spatial robustness of cross-view geo-localization retrieval ranking results.

Key words: Unmanned Aerial Vehicle (UAV) scenario, dynamic metric learning, cross-view geo-localization, image retrieval, deep learning

摘要:

现有跨视角地理定位研究主要聚焦于判断查询图像是否准确对应预定义图集中某个特定地理位置, 然而这种研究模式往往忽略了地理空间中固有的大量多空间尺度结构信息。为了实现更为稳健的定位效果, 模型不仅需要捕捉局部建筑细节, 还需理解通过建筑群和环境特征体现的目标之间的空间关系, 从而在不同空间尺度下提高定位准确性。为应对这些挑战, 提出多空间尺度跨视角地理定位任务, 并专门为此任务构建ML-Campus(Multi-Level Campus)数据集。ML-Campus数据集包含多视角、多来源的建筑图像, 并为每个图像标注了多空间尺度标签, 以体现不同空间尺度下的关联性和连续性。基于该数据集, 对现有的跨视角地理定位方法进行了实证评估, 以此为基准衡量其在此背景下的性能表现。为了进一步提升模型性能, 使用提出的CV-HAPPIER(Cross-View HAPPIER)方法进行训练, 以增强模型在不同空间尺度下的特征表示能力。大量在ML-Campus数据集上的实验结果表明, CV-HAPPIER方法显著提升了跨视角地理定位检索排名结果的空间鲁棒性。

关键词: 无人机场景, 动态度量学习, 跨视角地理定位, 图像检索, 深度学习