[1] ANTOL S, AGRAWAL A, LU J, et al. Vqa: Visual question answering[C]//Proceedings of the IEEE international conference on computer vision. 2015: 2425-2433.
[2] PARK S M, KIM Y G. Visual language integration: A survey and open challenges[J]. Computer Science Review, 2023, 48: 100548.
[3] MA L, LU Z, LI H. Learning to answer questions from image using convolutional neural network[C]//Proceedings of the AAAI conference on artificial intelligence. 2016, 30(1).
[4] CHEN K, WANG J, CHEN L C, et al. Abc-cnn: An attention based convolutional neural network for visual question answering[J]. arXiv preprint arXiv:1511.05960, 2015.
[5] ANDERSON P, HE X, BUEHLER C, et al. Bottom-up and top-down attention for image captioning and visual question answering[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 6077-6086.
[6] REN S, HE K, GIRSHICK R, et al. Faster R-CNN: Towards real-time object detection with region proposal networks[J]. IEEE transactions on pattern analysis and machine intelligence, 2016, 39(6): 1137-1149.
[7] YANG Z, HE X, GAO J, et al. Stacked attention networks for image question answering[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 21-29.
[8] YUAN W, QIU C. Cross-Modal Noise Information Elimination Network for Medical Visual Question Answering[C]//2024 4th International Conference on Neural Networks, Information and Communication Engineering (NNICE). IEEE, 2024: 813-819.
[9] PENG P, FAN W, SHEN Y, et al. A Global Visual Information Intervention Model for Medical Visual Question Answering[J]. Computers in Biology and Medicine, 2025, 192: 110195.
[10] PARK K R, LEE H J, KIM J U. Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality[C]//European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024: 42-59.
[11] WANG R, CHEN H, YANG J, et al. Adaptive sparse triple convolutional attention for enhanced visual question answering[J]. The Visual Computer, 2025: 1-17.
[12] MIAO J, YU K, FANG B, et al. Causality guided co-attention network for visual question answering[J]. Neural Networks, 2025: 108200.
[13] LU J, YANG J, BATRA D, et al. Hierarchical question-image co-attention for visual question answering[J]. Advances in neural information processing systems, 2016, 29.
[14] KIM J H, JUN J, ZHANG B T. Bilinear attention networks[J]. Advances in neural information processing systems, 2018, 31.
[15] YU Z, YU J, CUI Y, et al. Deep modular co-attention networks for visual question answering[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019: 6281-6290.
[16] YUSURF A A, FENG C, MAO X, et al. Multi-scale dual-stream visual feature extraction and graph reasoning for visual question answering[J]. Applied Intelligence, 2025, 55(6): 1-18.
[17] KE L, SHI X, WU H, et al. Text-image pair joint enhancement VQA model based on cross-modal attention[C]//Second International Conference on Image Processing and Artificial Intelligence (ICIPAI 2025). SPIE, 2025, 13780: 195-200.
[18] HE K, ZHANG X, REN S, et al. Deep residual learning for image recognition[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770-778.
[19] KRISHNA R, ZHU Y, GROTH O, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations[J]. International journal of computer vision, 2017, 123: 32-73.
[20] DEVLIN J, CHANG M W, LEE K, et al. Bert: Pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 2019: 4171-4186.
[21] HOCHREITER S, SCHMIDHUBER J. Long short-term memory[J]. Neural computation, 1997, 9(8): 1735-1780.
[22] 何世阳, 王朝晖, 龚声蓉, 钟珊. 基于跨模态信息过滤的视觉问答网络[J]. 计算机科学, 2024, 51(5): 85-91.HE Shiyang, WANG Zhaohui, GONG Shengrong, ZHONG Shan. Cross-modal Information Filtering based Networks for Visual Question Answering [J]. Computer Science, 2024, 51(5): 85-91.
[23] 姜丽梅, 李秉龙. 面向图像文本的多模态处理方法综述[J]. 计算机应用研究, 2024, 41(05):1281-1290. DOI:10.19734/j.issn.1001-3695.2023.08.0398. JIANG LI MEI, LI BING LONG. Comprehensive review of multimodal processing methods for image-text[J]. Application Research of Computers/Jisuanji Yingyong Yanjiu, 2024, 41(05):1281-1290. DOI:10. 19734/j. issn. 1001-3695. 2023. 08. 0398.
[24] 陈巧红, 项深祥, 方贤, 等. 跨模态自适应特征融合的视觉问答方法[J]. 哈尔滨工业大学学报, 2025,57(04): 94-104. CHEN QIAOHONG, XIANG SHENXIANG, FANG XIAN, et al. Visual question answering method based on cross-modal adaptive feature fusion[J]. Journal of Harbin Institute of Technology,2025,57(04):94-104.
[25] VASWANI A, SHAZEER N, PARMAR N, et al. Attention is all you need[J]. Advances in neural information processing systems, 2017, 30.
[26] HUDSON D A, MANNING C D. Gqa: A new dataset for real-world visual reasoning and compositional question answering[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019: 6700-6709.
[27] ASRI H S, SAFABAKHSH R. Advanced visual and textual co-context aware attention network with dependent multimodal fusion block for visual question answering[J]. Multimedia Tools and Applications, 2024, 83(40): 87959-87986.
[28] KOSHTI D, GUPTA A, KALLA M, et al. Trans-vqa: Fully transformer-based image question-answering model using question-guided vision attention[J]. Inteligencia Artificial, 2024, 27(73): 111-128.
[29] SHEN X, HAN D, ZONG L, et al. Relational reasoning and adaptive fusion for visual question answering: X. Shen et al[J]. Applied Intelligence, 2024, 54(6): 5062-5080.
[30] SHI J, HAN D, CHEN C, et al. SAFFNet: self-attention based on Fourier frequency domain filter network for visual question answering[J]. The Visual Computer, 2025, 41(8): 6149-6167.
|