| 1 |
HJELM R D, FEDOROV A, LAVOIE-MARCHILDON S, et al. Learning deep representations by mutual information estimation and maximization[EB/OL]. [2024-12-19]. https://arxiv.org/abs/1808.06670.
|
| 2 |
张丽英, 裴韬, 陈宜金, 等. 基于街景图像的城市环境评价研究综述. 地球信息科学学报, 2019, 21 (1): 46- 58.
|
|
ZHANG L Y , PEI T , CHEN Y J , et al. A review of urban environmental assessment based on street view images. Journal of Geo-Information Science, 2019, 21 (1): 46- 58.
|
| 3 |
DOERSCH C , SINGH S , GUPTA A , et al. What makes Paris look like Paris?. Communications of the ACM, 2015, 58 (12): 103- 110.
doi: 10.1145/2830541
|
| 4 |
NGUYEN Q C , SAJJADI M , MCCULLOUGH M , et al. Neighbourhood looking glass: 360° automated characterisation of the built environment for neighbourhood effects research. Journal of Epidemiology and Community Health, 2018, 72 (3): 260- 266.
doi: 10.1136/jech-2017-209456
|
| 5 |
ZHANG F , ZHANG D , LIU Y , et al. Representing place locales using scene elements. Computers, Environment and Urban Systems, 2018, 71, 153- 164.
doi: 10.1016/j.compenvurbsys.2018.05.005
|
| 6 |
DEWI C , CHEN R C , ZHUANG Y C , et al. Image enhancement method utilizing YOLO models to recognize road markings at night. IEEE Access, 2024, 12, 131065- 131081.
doi: 10.1109/ACCESS.2024.3440253
|
| 7 |
ZHAO H S, SHI J P, QI X J, et al. Pyramid scene parsing network[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu, USA: IEEE Press, 2017: 6230-6239.
|
| 8 |
LIANG X C , ZHAO T H , BILJECKI F . Revealing spatio-temporal evolution of urban visual environments with street view imagery. Landscape and Urban Planning, 2023, 237, 104802.
doi: 10.1016/j.landurbplan.2023.104802
|
| 9 |
XU S X, ZHANG C H, FAN L B, et al. AddressCLIP: empowering vision-language models for city-wide image address localization[C]//Proceedings of ECCV 2024. Berlin, Germany: Springer, 2025: 76-92.
|
| 10 |
NGIAM J, KHOSLA A, KIM M, et al. Multimodal deep learning[C]//Proceedings of the International Conference on Machine Learning (ICML). [S. l. ]: PMLR, 2011: 689-696.
|
| 11 |
HE W T , MA H J , LI S H , et al. Using augmented small multimodal models to guide large language models for multimodal relation extraction. Applied Sciences, 2023, 13 (22): 12208.
doi: 10.3390/app132212208
|
| 12 |
OUYANG T J, ZHANG X, HAN Z Y, et al. Health CLIP: depression rate prediction using health related features in satellite and street view images[C]//Proceedings of the ACM Web Conference 2024. New York, USA: ACM Press, 2024: 1142-1145.
|
| 13 |
|
| 14 |
LEE Y J , LI C Y , LIU H T , et al. Visual instruction tuning. Advances in Neural Information Processing Systems, 2023, 23, 34892- 34916.
|
| 15 |
ZHANG Y , ZHANG F , CHEN N C . Migratable urban street scene sensing method based on vision language pre-trained model. International Journal of Applied Earth Observation and Geoinformation, 2022, 113, 102989.
doi: 10.1016/j.jag.2022.102989
|
| 16 |
|
| 17 |
|
| 18 |
RADFORD A, KIM J W, HALLACY C, et al. Learning transferable visual models from natural language supervision[C]//Proceedings of International Conference on Machine Learning. [S. l. ]: PMLR, 2021: 8748-8763.
|
| 19 |
|
| 20 |
BAI Y , ZHAO Y , SHAO Y J , et al. Deep learning in different remote sensing image categories and applications: status and prospects. International Journal of Remote Sensing, 2022, 43 (5): 1800- 1847.
doi: 10.1080/01431161.2022.2048319
|
| 21 |
徐永智. 基于街景影像的建筑物底部轮廓提取[D]. 北京: 北京建筑大学, 2017.
|
|
XU Y Z. Extracting building footprints from digital measurable images[D]. Beijing: Beijing University of Civil Engineering and Architecture, 2017. (in Chinese)
|
| 22 |
LU J S, YANG J W, BATRA D, et al. Neural baby talk[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Salt Lake City, USA: IEEE Press, 2018: 7219-7228.
|
| 23 |
CAMPBELL A , BOTH A , SUN Q . Detecting and mapping traffic signs from Google Street View images using deep learning and GIS. Computers, Environment and Urban Systems, 2019, 77, 101350.
doi: 10.1016/j.compenvurbsys.2019.101350
|
| 24 |
QIU W S , ZHANG Z Y , LIU X , et al. Subjective or objective measures of street environment, which are more effective in explaining housing prices?. Landscape and Urban Planning, 2022, 221, 104358.
doi: 10.1016/j.landurbplan.2022.104358
|
| 25 |
|
| 26 |
SHIHAB I F, BHAGAT S R, SHARMA A. Precise and robust sidewalk detection: leveraging ensemble learning to surpass LLM limitations in urban environments[EB/OL]. [2024-12-19]. https://arxiv.org/abs/2405.14876.
|
| 27 |
REDMON J, FARHADI A. YOLO9000: better, faster, stronger[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu, USA: IEEE Press, 2017: 6517-6525.
|
| 28 |
CORDTS M, OMRAN M, RAMOS S, et al. The cityscapes dataset for semantic urban scene understanding[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas, USA: IEEE Press, 2016: 3213-3223.
|
| 29 |
ZHOU B L , ZHAO H , PUIG X , et al. Semantic understanding of scenes through the ADE20K dataset. International Journal of Computer Vision, 2019, 127 (3): 302- 321.
doi: 10.1007/s11263-018-1140-0
|
| 30 |
ZHANG Y , LIU P Y , BILJECKI F . Knowledge and topology: a two layer spatially dependent graph neural networks to identify urban functions with time-series street view image. ISPRS Journal of Photogrammetry and Remote Sensing, 2023, 198, 153- 168.
doi: 10.1016/j.isprsjprs.2023.03.008
|
| 31 |
WU M L , HUANG Q Y , GAO S , et al. Mixed land use measurement and mapping with street view images and spatial context-aware prompts via zero-shot multimodal learning. International Journal of Applied Earth Observation and Geoinformation, 2023, 125, 103591.
doi: 10.1016/j.jag.2023.103591
|
| 32 |
ZHAO Y H, ZHONG E H, YUAN C Y, et al. TG-LMM: enhancing medical image segmentation accuracy through text-guided large multi-modal model[EB/OL]. [2024-12-19]. https://arxiv.org/abs/2409.03412.
|
| 33 |
PICARD C, EDWARDS K M, DORIS A C, et al. From concept to manufacturing: evaluating vision-language models for engineering design[EB/OL]. [2024-12-19]. https://arxiv.org/abs/2311.12668.
|
| 34 |
YANG Y F , WANG S Q , LI D Y , et al. GeoLocator: a location-integrated Large Multimodal Model (LMM) for inferring geo-privacy. Applied Sciences, 2024, 14 (16): 7091.
doi: 10.3390/app14167091
|
| 35 |
JAYATI S, CHOI E, BURTON H, et al. Leveraging large multimodal models to augment image-based building damage assessment[C]//Proceedings of the 7th ACM SIGSPATIAL International Workshop on AI for Geographic Knowledge Discovery. New York, USA: ACM Press, 2024: 79-85.
|
| 36 |
HAO X X, CHEN W, YAN Y B, et al. UrbanVLP: multi-granularity vision-language pretraining for urban socioeconomic indicator prediction[EB/OL]. [2024-12-19]. https://arxiv.org/abs/2403.16831.
|
| 37 |
DOSOVITSKIY A, BEYER L, KOLESNIKOV A, et al. An image is worth 16×16 words: transformers for image recognition at scale[EB/OL]. [2024-12-19]. https://arxiv.org/abs/2010.11929.
|
| 38 |
|
| 39 |
TOLSTIKHIN I O , HOUISBY N , KOLESNIKOV A , et al. MLP-Mixer: an all-MLP architecture for vision. Advances in Neural Information Processing Systems, 2021, 34, 24261- 24272.
|
| 40 |
|
| 41 |
|
| 42 |
杨冬菊, 黄俊涛. 基于大语言模型的中文科技文献标注方法. 计算机工程, 2024, 50 (9): 113- 120.
doi: 10.19678/j.issn.1000-3428.0068400
|
|
YANG D J , HUANG J T . Chinese scientific literature annotation method based on large language model. Computer Engineering, 2024, 50 (9): 113- 120.
doi: 10.19678/j.issn.1000-3428.0068400
|
| 43 |
YASEEN M. What is YOLOv9: an in-depth exploration of the internal features of the next-generation object detector[EB/OL]. [2024-12-19]. https://arxiv.org/abs/2409.07813.
|
| 44 |
|
| 45 |
|
| 46 |
CHEN Z, WANG W Y, CAO Y, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling[EB/OL]. [2024-12-19]. https://arxiv.org/abs/2412.05271.
|
| 47 |
|
| 48 |
WANG P, BAI S, TAN S N, et al. Qwen2-VL: enhancing vision-language model's perception of the world at any resolution[EB/OL]. [2024-12-19]. https://arxiv.org/abs/2409.12191.
|
| 49 |
|
| 50 |
TEAM G, GEORGIEV P, LEI V I, et al. Gemini 1.5: unlocking multimodal understanding across millions of tokens of context[EB/OL]. [2024-12-19]. https://arxiv.org/abs/2403.05530.
|
| 51 |
CHIANG W L, ZHENG L M, SHENG Y, et al. Chatbot arena: an open platform for evaluating LLMs by human preference[EB/OL]. [2024-12-19]. https://arxiv.org/abs/2403.04132.
|