[1] SIMON S A. Value Judgments in Judicial Reasoning, and the Instability of the Fact-Law Distinction[J]. U. Cin. L. Rev., 2023, 92: 455.
[2] CASPER S, DAVIES X, SHI C, et al. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback[J/OL]. Transactions on Machine Learning Research, 2023.
[3] OUYANG L, WU J, JIANG X, et al. Training language models to follow instructions with human feedback[J]. Advances in neural information processing systems, 2022, 35: 27730-27744.
[4] 刘延飞, 李超, 王忠, 王杰铃. 多智能体深度强化学习及可扩展性研究进展[J]. 计算机工程与应用, 2025, 61(4): 1-24.LIU Y, LI C, WANG Z, WANG J L. Research Progress on Multi-Agent Deep Reinforcement Learning and Scalability[J]. Computer Engineering and Applications, 2025, 61(4): 1-24.
[5] BENCH-CAPON T, ARASZKIEWICZ M, ASHLEY K, et al. A history of AI and Law in 50 papers: 25 years of the international conference on AI and Law[J]. Artificial Intelligence and Law, 2012, 20(3): 215-319.
[6] SURDEN H. Artificial intelligence and law: An overview[J]. Ga. St. UL Rev., 2018, 35: 1305.
[7] HARÐARSON Þ H, LOFTSSON H, ÓLAFSSON S. Aligning language models for Icelandic legal text summarization[C]//Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025). 2025: 241-251.
[8] STIENNON N, OUYANG L, WU J, et al. Learning to summarize with human feedback[J]. Advances in neural information processing systems, 2020, 33: 3008-3021.
[9] LAMBERT N, PYATKIN V, MORRISON J, et al. Rewardbench: Evaluating reward models for language modeling[C]//Findings of the Association for Computational Linguistics: NAACL 2025. 2025: 1755-1797.
[10] DAHL M, MAGESH V, SUZGUN M, et al. Large legal fictions: Profiling legal hallucinations in large language models[J]. Journal of Legal Analysis, 2024, 16(1): 64-93.
[11] MCLAREN S, ROWE L. “You’re right to be skeptical!”: The Role of Legal Information Professionals in Assessing Generative AI Outputs[J]. Legal Information Management, 2025, 25(1): 19-25.
[12] GAO L, SCHULMAN J, HILTON J. Scaling laws for reward model overoptimization[C]//International Conference on Machine Learning. PMLR, 2023: 10835-10866.
[13] MINSKY M. Steps toward artificial intelligence[J]. Proceedings of the IRE, 1961, 49(1): 8-30.
[14] HOWARD R A. Dynamic programming and markov processes[M]. New York: John Wiley, 1960.
[15] BELLMAN R. Dynamic programming[J]. Science, 1966, 153(3731): 34-37.
[16] WATKINS C J C H. Learning from Delayed Rewards[D]. Cambridge: University of Cambridge, 1989.
[17] MNIH V, KAVUKCUOGLU K, SILVER D, et al. Playing atari with deep reinforcement learning[J]. arXiv:1312.5602, 2013.
[18] SILVER D, SCHRITTWIESER J, SIMONYAN K, et al. Mastering the game of Go without human knowledge[J]. Nature, 2017, 550: 354-359.
[19] BUCHANAN B G, HEADRICK T E. Some speculation about artificial intelligence and legal reasoning[J]. Stanford Law Review, 1970, 23: 40-62.
[20] HENDERSON P, KRASS M, ZHENG L, et al. Pile of law: Learning responsible data filtering from the law and a 256gb open-source legal dataset[J]. Advances in Neural Information Processing Systems, 2022, 35: 29217-29234.
[21] FEI Z, SHEN X, ZHU D, et al. Lawbench: Benchmarking legal knowledge of large language models[C]//Proceedings of the 2024 conference on empirical methods in natural language processing. 2024: 7933-7962.v
[22] BOMMARITO M, KATZ D M. GPT Takes the Bar Exam[J]. arXiv e-prints, 2022: arXiv: 2212.14402.
[23] ZHOU Z, YU K Y, TIAN S Y, et al. LawGPT: Knowledge-Guided Data Generation and Its Application to Legal LLM[J]. arXiv e-prints, 2025: arXiv: 2502.06572.
[24] GAO C, JIANG H, CAI D, et al. Strategyllm: Large language models as strategy generators, executors, optimizers, and evaluators for problem solving[J]. Advances in Neural Information Processing Systems, 2024, 37: 96797-96846.
[25] WANG Z, ZENG J, DELALLEAU O, et al. HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages[J]. arXiv e-prints, 2025: arXiv: 2505.11475.
[26] CHRISTIANO P F, LEIKE J, BROWN T, et al. Deep reinforcement learning from human preferences[J]. Advances in neural information processing systems, 2017, 30.
[27] PENG B, LI C, HE P, et al. Instruction Tuning with GPT-4[J]. arXiv e-prints, 2023: arXiv: 2304.03277.
[28] MIAO Y, ZHANG S, DING L, et al. Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling[J]. Advances in Neural Information Processing Systems, 2024, 37: 134387-134429.
[29] ZHANG M J, WANG Z, HWANG J D, et al. Diverging Preferences: When do Annotators Disagree and do Models Know?[C]//International Conference on Machine Learning. PMLR, 2025: 76193-76212.
[30] RAFAILOV R, SHARMA A, MITCHELL E, et al. Direct preference optimization: Your language model is secretly a reward model[J]. Advances in neural information processing systems, 2023, 36: 53728-53741.
[31] SHAO Z, WANG P, ZHU Q, et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models[J]. arXiv e-prints, 2024: arXiv: 2402.03300.
[32] SNELL C, KOSTRIKOV I, SU Y, et al. Offline RL for Natural Language Generation with Implicit Language Q Learning[J]. arXiv e-prints, 2022: arXiv: 2206.11871.
[33] DENG R, FENG D, LEI W. AMaPO: Adaptive Margin-attached Preference Optimization for Language Model Alignment[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2026, 40(44): 37341-37349.
[34] LYU J, MA X, LI X, et al. Mildly conservative q-learning for offline reinforcement learning[J]. Advances in Neural Information Processing Systems, 2022, 35: 1711-1724.
[35] TARASOV D, KURENKOV V, NIKULIN A, et al. Revisiting the minimalist approach to offline reinforcement learning[J]. Advances in Neural Information Processing Systems, 2023, 36: 11592-11620.
[36] ZHANG H, GUI L, LEI Y, et al. Copr: Continual human preference learning via optimal policy regularization[C]//Findings of the Association for Computational Linguistics: ACL 2025. 2025: 5377-5398.
[37] CAI H, ZHAO S, ZHANG L, et al. Unilaw-r1: A large language model for legal reasoning with reinforcement learning and iterative inference[C]//Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025: 18128-18142.
[38] NATH A, GRAFF C, BACHININ A, et al. Frictional Agent Alignment Framework: Slow Down and Don’t Break Things[C]//Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025: 11042-11089.
[39] LAI J, HE K, TAN Y, et al. Large Language Models in Law: A Survey[J]. International Journal of Law and Information Technology, 2024, 32(1): 1-32.
[40] LE H, TRAN Q H, NGUYEN D, et al. Multi-reference preference optimization for large language models[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2025, 39(23): 24375-24383.
[41] ZHOU J, JI J, DAI J, et al. Sequence to sequence reward modeling: Improving rlhf by language feedback[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2025, 39(26): 27765-27773.
[42] NIGAM S K, TYAGI T, SHUKLA S, et al. ReGal: A First Look at PPO-based Legal AI for Judgment Prediction and Summarization in India[C]//Bridge between Artificial Intelligence and Law. 2025.
[43] LEE H, PHATALE S, MANSOOR H, et al. Rlaif: Scaling reinforcement learning from human feedback with ai feedback[J]. 2023.
[44] FÜRST J. The Experts know it all: Reinforcement Learning from Human Feedback for Legal Information Extraction[J]. 2025.
[45] LIN Q, LU H, YUAN C, et al. Data with high and consistent preference difference are better for reward model[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2025, 39(26): 27482-27490.
[46] LIU S, SHEN X, LAI Y, et al. Haf-rm: A hybrid alignment framework for reward model training[C]//Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025: 18874-18893.
[47] LI Y C, XU T, YU Y, et al. Generalist reward models: Found inside large language models[J]. arXiv preprint arXiv:2506.23235, 2025.
[48] DECHTIAR M, KATZ D M, SUNDARESAN M, et al. GRAPH-GRPO-LEX: Contract Graph Modeling and Reinforcement Learning with Group Relative Policy Optimization[C]//2025 IEEE International Conference on Data Mining Workshops (ICDMW). IEEE, 2025: 767-776.
[49] ZHONG H, WANG Y, TU C, et al. Iteratively questioning and answering for interpretable legal judgment prediction[C]//Proceedings of the AAAI conference on artificial intelligence. 2020, 34(01): 1250-1257.
[50] WU Y, WANG C, GUMUSEL E, et al. Knowledge-infused legal wisdom: Navigating llm consultation through the lens of diagnostics and positive-unlabeled reinforcement learning[C]//Findings of the Association for Computational Linguistics: ACL 2024. 2024: 15542-15555.
[51] ZHOU Y, ZANETTE A. ArCHer: training language model agents via hierarchical multi-turn RL[C]//Proceedings of the 41st International Conference on Machine Learning. 2024: 62178-62209.
[52] ZHONG W, CUI R, GUO Y, et al. Agieval: A human-centric benchmark for evaluating foundation models[C]//Findings of the association for computational linguistics: NAACL 2024. 2024: 2299-2314.
[53] ZANATTI M, RIBEIRO R, PINTO H S. Exploring Metric Correlations for Legal Text Summarization Evaluation[C]//Proceedings of the Twentieth International Conference on Artificial Intelligence and Law. 2025: 389-393.
[54] HE C, HU H, LI Y, et al. A survey of large language models for legal tasks: Progress, prospects and challenges[J]. Computer Science Review, 2026, 60: 100906.
[55] YU F, SEEDAT N, HERRMANNOVA D, et al. Beyond pointwise scores: Decomposed criteria-based evaluation of llm responses[C]//Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track. 2025: 1931-1954.
[56] GUHA N, NYARKO J, HO D, et al. Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models[J]. Advances in neural information processing systems, 2023, 36: 44123-44279.
[57] LI H, CHEN Y, AI Q, et al. Lexeval: A comprehensive chinese legal benchmark for evaluating large language models[J]. Advances in Neural Information Processing Systems, 2024, 37: 25061-25094.
[58] ZHENG L, CHIANG W L, SHENG Y, et al. Judging llm-as-a-judge with mt-bench and chatbot arena[J]. Advances in neural information processing systems, 2023, 36: 46595-46623.
[59] ADAK S, CHATTERJEE P, BANERJEE S, et al. AURA: Affordance-Understanding and Risk-aware Alignment Technique for Large Language Models[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2026, 40(44): 37204-37212.
[60] LAI Y, WANG S, LIU S, et al. Alarm: Align language models via hierarchical rewards modeling[C]//Findings of the Association for Computational Linguistics: ACL 2024. 2024: 7817-7831.
[61] AHMADIAN A, CREMER C, GALLÉ M, et al. Back to basics: Revisiting REINFORCE-style optimization for learning from human feedback in LLMs[C]//Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024: 12248-12267.
[62] ZHANG M J, WANG Z, HWANG J D, et al. Diverging Preferences: When do Annotators Disagree and do Models Know?[C]//International Conference on Machine Learning. PMLR, 2025: 76193-76212.
[63] AL-YAYCHLI B A, MOHAMMED N F, SAFAR M, et al. Automating rlhf-based hallucination tracking[J]. Journal Européen des Systèmes Automatisés, 2025, 58(7): 1477.
[64] HUANG L, YU W, MA W, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions[J]. ACM Transactions on Information Systems, 2025, 43(2): 1-55.
[65] VAN DIJCK G, AGUILERA C, CHAKRAVARTHY S M. Deciphering disagreement in the annotation of EU legislation[J]. Artificial Intelligence and Law, 2024: 1-36.
[66] 最高人民法院关于规范和加强人工智能司法应用的意见[N].人民法院报,2022-12-10(004).DOI:10.28650/n.cnki.nrmfy.2022.004314.Opinions of the Supreme People's Court on Regulating and Strengthening the Application of Artificial Intelligence in Judicial Affairs [N]. People's Court Daily, 2022-12-10 (004). DOI: 10.28650/n.cnki.nrmfy.2022.004314.
[67] DAHLGREN LINDSTRÖM A, METHNANI L, KRAUSE L, et al. Helpful, harmless, honest? Sociotechnical limits of AI alignment and safety through Reinforcement Learning from Human Feedback: AD Lindström et al[J]. Ethics and Information Technology, 2025, 27(2): 28.
[68] LI Z, FANG F, ZHANG X, et al. Not All Voices Are Rewarded Equally: Probing and Repairing Reward Models across Human Diversity[C]//Findings of the Association for Computational Linguistics: EMNLP 2025. 2025: 3426-3455.
[69] LOPES G P. Bias in adjudication and the promise of AI: Challenges to procedural fairness[J]. Law, Technology and Humans, 2025, 7(1): 47-67.
[70] 张凌寒,于琳.生成式人工智能价值链上的侵权责任划分[J].北京行政学院学报,2025,(05):72-83.ZHANG LH, YU L. Division of Liability for Infringement in the Value Chain of Generative Artificial Intelligence [J]. Journal of Beijing Administration Institute, 2025, (0
5): 72-83.
[71] 吴汉东,樊赛尔.试论生成式人工智能服务提供者的合理注意义务[J].知识产权,2025,(11):3-23.WU HD, FAN SE. On the Reasonable Duty of Care of Generative Artificial Intelligence Service Providers [J]. Intellectual Property, 2025, (11): 3-23.]
[72] 生成式人工智能服务管理暂行办法[J].中华人民共和国国务院公报,2023,(24):39-42.Interim Measures for the Administration of Generative Artificial Intelligence Services [J]. Gazette of the State Council of the People's Republic of China, 2023, (24):39-42.
[73] 全国网络安全标准化技术委员会(SAC/TC 260).网络安全技术 生成式人工智能预训练和优化训练数据安全规范:GB/T 45652-2025[S].中国标准出版社,2025.National Technical Committee for Standardization of Cybersecurity (SAC/TC 260). Cybersecurity Technology—Security Specifications for Pre-training and Fine-tuning Data of Generative Artificial Intelligence: GB/T 45652-2025 [S]. China Standards Press, 2025.
[74] 洪涛.正当程序视角下人工智能辅助量刑的挑战与重塑[J].东北大学学报(社会科学版),2025,27(02):111-120.HONG T. Challenges and Reconstruction of Artificial Intelligence-Assisted Sentencing from the Perspective of Due Process [J]. Journal of Northeastern University (Social Sciences Edition), 2025, 27(02): 111-120.
[75] ETHAYARAJH K, CHOI Y, SWAYAMDIPTA S. Understanding dataset difficulty with $\mathcal {V} $-usable information[C]//International Conference on Machine Learning. PMLR, 2022: 5988-6008.
[76] ZHANG G, DUAN J. Vickreyfeedback: Cost-efficient data construction for reinforcement learning from human feedback[C]//International Conference on Principles and Practice of Multi-Agent Systems. Cham: Springer Nature Switzerland, 2024: 351-366.
[77] WANG H, XIONG W, XIE T, et al. Interpretable preferences via multi-objective reward modeling and mixture-of-experts[C]//Findings of the Association for Computational Linguistics: EMNLP 2024. 2024: 10582-10592.
[78] HU E J, SHEN Y, WALLIS P, et al. Lora: Low-rank adaptation of large language models[J]. Iclr, 2022, 1(2): 3.
[79] BANERJEE S, LAYEK S, TRIPATHY S, et al. Safeinfer: Context adaptive decoding time safety alignment for large language models[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2025, 39(26): 27188-27196.
[80] XU H, WU L, CHENG C, et al. Multi-Value Alignment for LLMs via Value Decorrelation and Extrapolation[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2026, 40(40): 34133-34141.
|