[1] 刘成林, 金连文, 白翔,等. 2023. 文档智能分析与识别前沿:回顾与展望. 中国图象图形学报, 28(08):2223-2252 DOI: 10.11834/jig.221112.
Liu Chenglin, Jin Lianwen, Bai Xiang, et al. 2023. Frontiers of intelligent document analysis and recognition: review and prospects. Journal of Image and Graphics, 28(08):2223-2252 DOI: 10.11834/jig.221112.
[2] 林泽柠, 汪嘉鹏, 金连文. 2023. 视觉信息抽取的深度学习方法综述. 中国图象图形学报, 28(08):2276-2297 DOI: 10.11834/jig.220904.
Lin Zening, Wang Jiapeng, Jin Lianwen. 2023. Visual information extraction deep learning method: a critical review. Journal of Image and Graphics, 28(08):2276-2297 DOI: 10.11834/jig.220904.
[3] 朱贵德, 黄海. 文本视觉问答综述[J]. 计算机工程, 2024, 50(2): 1-14.
Guide ZHU, Hai HUANG. Survey of Text-based Visual Question Answering[J]. Computer Engineering, 2024, 50(2): 1-14.
[4] 孙仁科, 许靖昊, 皇甫志宇, 李仲年, 许新征. 基于视觉-语言预训练模型的零样本迁移学习方法综述[J]. 计算机工程, 2024, 50(10): 1-15.
SUN Renke, XU Jinghao, HUANGFU Zhiyu, LI Zhongnian, XU Xinzheng. Survey of Zero-Shot Transfer Learning Methods Based on Vision-Language Pre-Trained Models[J]. Computer Engineering, 2024, 50(10): 1-15.
[5] Liu H, Li C, Wu Q, Lee Y J. Visual instruction tuning[J]. Advances in Neural Information Processing Systems, 2023, 36: 34892–34916.
[6] Chen Z, Wu J, Wang W, et al. InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2024: 24185–24198.
[7] Xu Y, Li M, Cui L, et al. LayoutLM: Pre-training of text and layout for document image understanding[C]// Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. New York: ACM, 2020: 1192–1200.
[8] Kim G, Hong T, Yim M, et al. OCR-free document understanding transformer[C]// European Conference on Computer Vision. Heidelberg: Springer, 2022: 498–517.
[9] DOSOVITSKIY A, BEYER L, KOLESNIKOV A, et al. An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale[C]//International Conference on Learning Representations (ICLR), 2021.
[10] Xu Y, Xu Y, Lv T, et al. LayoutLMv2: Multi-modal pre-training for visually-rich document understanding[C]// Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Stroudsburg: ACL, 2021: 2579–2591.
[11] Huang Y, Lv T, Cui L, et al. LayoutLMv3: Pre-training for document AI with unified text and image masking[C]// Proceedings of the 30th ACM International Conference on Multimedia. New York: ACM, 2022: 4083–4091.
[12] Liao M, Wan Z, Yao C, et al. Real-time scene text detection with differentiable binarization[C]// Proceedings of the AAAI Conference on Artificial Intelligence. Palo Alto: AAAI Press, 2020: 11474–11481.
[13] Gu J, Kuen J, Morariu V I, et al. Unidoc: Unified pretraining framework for document understanding[J]. Advances in Neural Information Processing Systems, 2021, 34: 39–50.
[14] Wang D, Raman N, Sibue M, et al. DocLLM: A layout-aware generative language model for multimodal document understanding[C]// Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Stroudsburg: ACL, 2024: 8529–8548.
[15] Hudson D A, Manning C D. GQA: A new dataset for real-world visual reasoning and compositional question answering[C]// Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2019: 6700–6709.
[16] Mathew M, Karatzas D, Jawahar C V. DocVQA: A dataset for VQA on document images[C]// Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. Piscataway: IEEE, 2021: 2200–2209.
[17] Xu Y, Lv T, Cui L, et al. XFUND: A benchmark dataset for multilingual visually rich form understanding[C]// Findings of the Association for Computational Linguistics: ACL 2022. Stroudsburg: ACL, 2022: 3214–3224.
[18] Sun Y, Ni Z, Chng C-K, et al. ICDAR 2019 competition on large-scale street view text with partial labeling—RRC-LSVT[C]// 2019 International Conference on Document Analysis and Recognition (ICDAR). Piscataway: IEEE, 2019: 1557–1562.
[19] Cui C, Zhang Y, Sun T, et al. PP-OCRv5: A specialized 5M-parameter model rivaling billion-parameter vision-language models on OCR tasks[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Piscataway: IEEE, 2026: 2467-2476.
[20] Bai S, Cai Y, Chen R, et al. Qwen3-VL technical report[J]. arXiv preprint arXiv:2511.21631, 2025.
[21] Lee K, Joshi M, Turc I R, et al. Pix2struct: Screenshot parsing as pretraining for visual language understanding[C]// International Conference on Machine Learning. Proceedings of Machine Learning Research (PMLR), 2023: 18893–18912.
[22] Bai J, Bai S, Yang S, et al. Qwen-VL: A versatile vision-language model for understanding, localization, text reading, and beyond[J]. arXiv preprint arXiv:2308.12966, 2023.
[23] Chen K, Zhang Z, Zeng W, et al. Shikra: Unleashing multimodal LLM's referential dialogue magic[J]. arXiv preprint arXiv:2306.15195, 2023.
[24] Dai W, Li J, Li D, et al. InstructBLIP: Towards general-purpose vision-language models with instruction tuning[J]. Advances in Neural Information Processing Systems, 2023, 36: 49250–49267.
[25] Lu J, Yu H, Wang Y, et al. A Bounding Box is Worth One Token-Interleaving Layout and Text in a Large Language Model for Document Understanding[C]//Findings of the Association for Computational Linguistics: ACL 2025. 2025: 7252-7273.
[26] Zhou Y, Chen Y, Lin H, et al. DOGR: Towards versatile visual document grounding and referring[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. 2025: 3596-3606.
|