[1] ACHIAM J, ADLER S, AGARWAL S, et al. Gpt-4 technical report[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2303.08774. [2] 李博, 季佰军, 段湘煜. 基于译文易错词纠正机制的大语言模型机器翻译[J]. 计算机工程, 2026, 52(2): 372-382. LI B, JI B J, DUAN X Y. Machine translation with large language models based on correction mechanism of error-prone words in translations[J]. Computer Engineering, 2026, 52(2): 372-382. (in Chinese) [3] XI Z H, CHEN W X, GUO X, et al. The rise and potential of large language model based agents: a survey[J]. Science China Information Sciences, 2025, 68(2): 121101. [4] 罗焕坤, 葛一烽, 刘帅. 大语言模型在数学推理中的研究进展[J]. 计算机工程, 2024, 50(9): 1-17. LUO H K, GE Y F, LIU S. Research progress of large language models in mathematical reasoning[J]. Computer Engineering, 2024, 50(9): 1-17. (in Chinese) [5] ZHANG X L, TIAN C X, YANG X J, et al. AlpaCare: instruction-tuned large language models for medical application[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2310.14558. [6] 沈晨晨, 岳圣斌, 刘书隽, 等. 面向法律领域的大模型微调与应用[J]. 大数据, 2024, 10(5): 11-27. SHEN C C, YUE S B, LIU S J, et al. Fine-tuning and application of large language model in law domain[J]. Big Data Research, 2024, 10(5): 11-27. (in Chinese) [7] JI Z W, LEE N, FRIESKE R, et al. Survey of hallucination in natural language generation[J]. ACM Computing Surveys, 2023, 55(12): 1-38. [8] ZHANG Y, LI Y F, CUI L Y, et al. Siren’s song in the AI ocean: a survey on hallucination in large language models[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2309.01219. [9] RAWTE V, SHETH A, DAS A. A survey of hallucination in large foundation models[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2309.05922. [10] HUANG L, YU W, MA W, et al. A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions[J]. ACM Transactions on Information Systems, 2025, 43(2): 1-55. [11] 刘泽垣, 王鹏江, 宋晓斌, 等. 大语言模型的幻觉问题研究综述[J]. 软件学报, 2025, 36(3): 1152-1185. LIU Z Y, WANG P J, SONG X B, et al. Survey on hallucinations in large language models[J]. Journal of Software, 2025, 36(3): 1152-1185. (in Chinese) [12] LINDLEY D V. On a measure of the information provided by an experiment[J]. The Annals of Mathematical Statistics, 1956, 27(4): 986-1005. [13] FARQUHAR S, KOSSEN J, KUHN L, et al. Detecting hallucinations in large language models using semantic entropy[J]. Nature, 2024, 630(8017): 625-630. [14] THORNE J, VLACHOS A, CHRISTODOULOPOULOS C, et al. FEVER: a large-scale dataset for fact extraction and VERification[C]//Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1(Long Papers). New Orleans, Louisiana: Association for Computational Linguistics, 2018: 809-819. [15] LIN S, HILTON J, EVANS O. TruthfulQA: measuring how models mimic human falsehoods[C]//Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Dublin, Ireland: Association for Computational Linguistics, 2022: 3214-3252. [16] HENDRYCKS D, BURNS C, BASART S, et al. Measuring massive multitask language understanding[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2009.03300. [17] DU Y L, LI S, TORRALBA A, et al. Improving factuality and reasoning in language models through multiagent debate[C]//Proceedings of International Conference on Machine Learning. [S. l.]: PMLR, 2024: 11733-11763. [18] TONMOY S M T I, MEHEDI ZAMAN S M, JAIN V, et al. A comprehensive survey of hallucination mitigation techniques in large language models[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2401.01313. [19] WHITE J, FU Q C, HAYS S, et al. A prompt pattern catalog to enhance prompt engineering with ChatGPT[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2302.11382. [20] SHUSTER K, POFF S, CHEN M Y, et al. Retrieval augmentation reduces hallucination in conversation[C]//Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2021. Punta Cana, Dominican Republic: Association for Computational Linguistics, 2021: 3784-3803. [21] DHULIAWALA S, KOMEILI M, XU J, et al. Chain-of-verification reduces hallucination in large language models[C]//Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024. Bangkok, Thailand: Association for Computational Linguistics, 2024: 3563-3578. [22] SHI W J, HAN X C, LEWIS M, et al. Trusting your evidence: hallucinate less with context-aware decoding[C]//Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers). Mexico City, Mexico: Association for Computational Linguistics, 2024: 783-791. [23] FATAHI B F, QIAN K, HAN B, et al. FLEEK: factual error detection and correction with evidence retrieved from external knowledge[C]//Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Singapore: Association for Computational Linguistics, 2023: 124-130. [24] CHRISTIANO P F, LEIKE J, BROWN T B, et al. Deep reinforcement learning from human preferences[C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. New York, USA: ACM Press, 2017: 4302-4310. [25] BRADLEY R A, TERRY M E. Rank analysis of incomplete block designs: I. the method of paired comparisons[J]. Biometrika, 1952, 39(3/4): 324-345. [26] SCHULMAN J, WOLSKI F, DHARIWAL P, et al. Proximal policy optimization algorithms[EB/OL]. [2025-01-02]. https://arxiv.org/abs/1707.06347. [27] OUYANG L, WU J, ALMEIDA D, et al. Training language models to follow instructions with human feedback[C]//Proceedings of NeurIPS 2022. New Orleans, USA: Neural Information Processing Systems Foundation, Inc., 2022: 27730-27744. [28] SCHULMAN J, MORITZ P, LEVINE S, et al. High-dimensional continuous control using generalized advantage estimation[EB/OL]. [2025-01-02]. https://arxiv.org/abs/1506.02438. [29] SUTTON R S. Learning to predict by the methods of temporal differences[J]. Machine Learning, 1988, 3(1): 9-44. [30] MALON C. Team Papelo: Transformer networks at FEVER[C]//Proceedings of the 1st Workshop on Fact Extraction and VERification (FEVER). Brussels, Belgium: Association for Computational Linguistics, 2018: 109-113. [31] VASWANI A, SHAZER N, PARMAR N, et al. Attention is all you need[C]//Proceedings of the 31st International Conference on Neural Information Processing Systems. New York, USA: ACM Press, 2017:5998-6008. [32] LIN C Y. ROUGE: a package for automatic evaluation of summaries[C]//Proceedings of the Workshop on Text Summarization Branches Out. Barcelona, Spain: Association for Computational Linguistics, 2004: 74-81. [33] PAPINENI K, ROUKOS S, WARD T, et al. BLEU: a method for automatic evaluation of machine translation[C]//Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. Philadelphia, USA: Association for Computational Linguistics, 2002: 311-318. [34] ZHANG T Y, KISHORE V, WU F, et al. BERTScore: evaluating text generation with BERT[EB/OL]. [2025-01-02]. https://arxiv.org/abs/1904.09675. [35] LI J W, GALLEY M, BROCKETT C, et al. A diversity-promoting objective function for neural conversation models[C]//Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. San Diego, USA: Association for Computational Linguistics, 2016: 110-119. [36] YOFFE L, AMAYUELAS A, WANG W Y. DebUnc: improving large language model agent communication with uncertainty metrics[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2407.06426. [37] WANG X Z, WEI J, SCHUURMANS D, et al. Self-consistency improves chain of thought reasoning in language models[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2203.11171. [38] BOSMA M, CHI E, ICHTER B, et al. Chain-of-thought prompting elicits reasoning in large language models[C]//Proceedings of NeurIPS 2022. New Orleans, USA: Neural Information Processing Systems Foundation, Inc., 2022: 24824-24837. [39] CASSANO F, GOPINATH A, NARASIMHAN K, et al. Reflexion: language agents with verbal reinforcement learning[C]//Proceedings of NeurIPS 2023. New Orleans, USA: Neural Information Processing Systems Foundation, Inc., 2023: 8634-8652. [40] CHUANG Y, XIE Y, LUO H, et al. DoLa: decoding by contrasting layers improves factuality in large language models[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2309.03883. [41] WANG A T, SONG L F, PENG B L, et al. Fine-grained self-endorsement improves factuality and reasoning[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2402.15631. [42] ALMAZROUEI E, ALOBEIDLI H, CAPPELLI A, et al. The RefinedWeb dataset for falcon LLM: outperforming curated corpora with Web data only[C]//Proceedings of NeurIPS 2023. New Orleans, Louisiana, USA: Neural Information Processing Systems Foundation, Inc., 2023: 79155-79172. [43] GLM T, ZENG A, XU B, et al. ChatGLM: a family of large language models from GLM-130B to GLM-4 all tools[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2406.12793. [44] BAI J Z, BAI S, CHU Y F, et al. Qwen technical report[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2309.16609. [45] CUI Y M, YANG Z Q, YAO X. Efficient and effective text encoding for Chinese LLaMA and alpaca[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2304.08177. [46] TOUVRON H, MARTIN L, STONE K, et al. Llama 2: open foundation and fine-tuned chat models[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2307.09288. [47] HU E J, SHEN Y L, WALLIS P, et al. LoRA: low-rank adaptation of large language models[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2106.09685. [48] WILLIAMS A, NANGIA N, BOWMAN S. A broad-coverage challenge corpus for sentence understanding through inference[C]//Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1(Long Papers). New Orleans, Louisiana: Association for Computational Linguistics, 2018: 1112-1122. [49] HE P C, LIU X D, GAO J F, et al. DeBERTa: decoding-enhanced BERT with disentangled attention[EB/OL]. [2025-01-02]. https://arxiv.org/abs/2006.03654. |