In plain words: Chatbot training has three stages: learning from raw text, copying human-written answers, and adjusting to human ratings. Reframing them as a two-player game between a responder and a judge explains why matching human values is hard and suggests new training ideas.
Abstract
By formally defining the training processes of large language models (LLMs), which usually encompasses pre-training, supervised fine-tuning, and reinforcement learning with human feedback, within a single and unified machine learning paradigm, we can glean pivotal insights for advancing LLM technologies. This position paper delineates the parallels between the training methods of LLMs and the strategies employed for the development of agents in two-player games, as studied in game theory, reinforcement learning, and multi-agent systems. We propose a re-conceptualization of LLM learning processes in terms of agent learning in language-based games. This framework unveils innovative perspectives on the successes and challenges in LLM development, offering a fresh understanding of addressing alignment issues among other strategic considerations. Furthermore, our two-player game approach sheds light on novel data preparation and machine learning techniques for training LLMs.
Yang Liu, Peng Sun, Hang Li
arXiv:2402.08078 · cs.CL, cs.LG · submitted Feb 12, 2024
abstract · pdf · html