[1] 曾碧卿, 陈鹏飞, 姚勇涛. 融合思维链和低秩自适应微调的方面情感三元组抽取[J]. 计算机工程, 2024, 50(7): 53-62.
[2] 顾滢双, 桂韬, 张奇. 基于语义熵反馈强化学习的大语言模型事实性幻觉缓解[J]. 计算机工程, doi: 10.19678/j.issn.1000-3428.0252253.
[3] Wu J, Gan W, Chen Z, et al. Multimodal large language models: A survey[C]//2023 IEEE International Conference on Big Data (BigData). IEEE, 2023: 2247-2256.
[4] Liu H, Xue W, Chen Y, et al. A survey on hallucination in large vision-language models[J]. arXiv preprint arXiv:2402.00253, 2024.
[5] Chen L, Zaharia M, Zou J. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance[J]. Transactions on Machine Learning Research.
[6] Team G, Anil R, Borgeaud S, et al. Gemini: a family of highly capable multimodal models[J]. arXiv preprint arXiv:2312.11805, 2023.
[7] Li L H, Yatskar M, Yin D, et al. Visualbert: A simple and performant baseline for vision and language[J]. arXiv preprint arXiv:1908.03557, 2019.
[8] Wang T, Zhang J, Fei J, et al. Caption anything: Interactive image description with diverse multimodal controls[J]. arXiv preprint arXiv:2305.02677, 2023.
[9] Leng S, Zhang H, Chen G, et al. Mitigating object hallucinations in large vision-language models through visual contrastive decoding[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024: 13872-13882.
[10] Rohrbach A, Hendricks L A, Burns K, et al. Object Hallucination in Image Captioning[C]//Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018: 4035-4045.
[11] Sun Z, Shen S, Cao S, et al. Aligning large multimodal models with factually augmented rlhf[C]//Findings of the Association for Computational Linguistics: ACL 2024. 2024: 13088-13110.
[12] Zhao Z, Wang B, Ouyang L, et al. Beyond hallucinations: Enhancing lvlms through hallucination-aware direct preference optimization[J]. arXiv preprint arXiv:2311.16839, 2023.
[13] Huang Q, Dong X, Zhang P, et al. Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024: 13418-13427.
[14] Liu S, Zheng K, Chen W. Paying more attention to image: A training-free method for alleviating hallucination in lvlms[C]//European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024: 125-140.
[15] Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C]//International conference on machine learning. PmLR, 2021: 8748-8763.
[16] Li J, Li D, Xiong C, et al. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation[C]//International conference on machine learning. PMLR, 2022: 12888-12900.
[17] Liu H, Li C, Wu Q, et al. Visual instruction tuning[J]. Advances in neural information processing systems, 2023, 36: 34892-34916.
[18] Zhu D, Chen J, Shen X, et al. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models[C]//The Twelfth International Conference on Learning Representations.
[19] Lu H, Liu W, Zhang B, et al. DeepseekVL: towards real-world vision-language understanding[J]. arXiv preprint arXiv:2403.05525, 2024.
[20] Chen Z, Wu J, Wang W, et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2024: 24185-24198.
[21] Jiang C, Xu H, Dong M, et al. Hallucination augmented contrastive learning for multimodal large language model[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024: 27036-27046.
[22] Jawahar G, Sagot B, Seddah D. What does BERT learn about the structure of language?[C]//ACL 2019-57th Annual Meeting of the Association for Computational Linguistics. 2019.
[23] Tenney I, Das D, Pavlick E. BERT Rediscovers the Classical NLP Pipeline[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019: 4593-4601.
[24] Liang V W, Zhang Y, Kwon Y, et al. Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning[J]. Advances in Neural Information Processing Systems, 2022, 35: 17612-17625.
[25] Li Y, Du Y, Zhou K, et al. Evaluating Object Hallucination in Large Vision-Language Models[C]//Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023: 292-305. POPE
[26] Fu C, Chen P, Shen Y, et al. Mme: A comprehensive evaluation benchmark for multimodal large language models[C]//The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track. 2025.
[27] An W, Tian F, Leng S, et al. Mitigating object hallucinations in large vision-language models with assembly of global and local attention[C]//Proceedings of the Computer Vision and Pattern Recognition Conference. 2025: 29915-29926.
[28] Anand N, Jha S, Bamba U, et al. CRoPS: A Training-Free Hallucination Mitigation Framework for Vision-Language Models[J]. Transactions on Machine Learning Research.
[29] Tang F, Liu C, Xu Z, et al. Seeing far and clearly: Mitigating hallucinations in mllms with attention causal decoding[C]//Proceedings of the Computer Vision and Pattern Recognition Conference. 2025: 26147-26159.
|