In plain words: A language model learns to search and read pages in a text-only browser, writing answers with source links so readers can verify them. After copying human browsing and refining with human preference judgments, its answers were preferred over human demonstrators' 56% of the time.
Abstract
We fine-tune GPT-3 to answer long-form questions using a text-based web-browsing environment, which allows the model to search and navigate the web. By setting up the task so that it can be performed by humans, we are able to train models on the task using imitation learning, and then optimize answer quality with human feedback. To make human evaluation of factual accuracy easier, models must collect references while browsing in support of their answers. We train and evaluate our models on ELI5, a dataset of questions asked by Reddit users. Our best model is obtained by fine-tuning GPT-3 using behavior cloning, and then performing rejection sampling against a reward model trained to predict human preferences. This model's answers are preferred by humans 56% of the time to those of our human demonstrators, and 69% of the time to the highest-voted answer from Reddit.
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, et al.
arXiv:2112.09332 · cs.CL, cs.AI, cs.LG · submitted Dec 17, 2021 · updated Jun 1, 2022
abstract · pdf · html · 32 pages
After that the next step I think is to make LMs that map arguments to functions from unstructured input, externally run the function and use the result in the next token predictions. LMs could write functions and execute them as regular code. They could learn to generate task descriptions, code, constraints and tests, and then run them in the loop to get experimental feedback. Regular code could call the neural net, and the neural net could call regular code. This would bring neural nets closer to symbolic AI.
So I see LMs as developers, iterating towards solutions. They have access to efficient simulation and compilers to augment the neural part.
[1] "Improving language models by retrieving from trillions of tokens" https://arxiv.org/pdf/2112.04426.pdf
[2] "Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents" https://arxiv.org/pdf/2201.07207.pdf
[3] "Pretrained Transformers As Universal Computation Engines" https://arxiv.org/pdf/2103.05247.pdf