In plain words: A language model writes a Python program over data structures holding natural-language knowledge, and an interpreter runs it to give the answer. With one general prompt, it beat strong baselines on math, classification, question answering, and instruction following, and the code shows the reasoning.
Abstract · Natural Language Embedded Programs for Hybrid Language Symbolic Reasoning
How can we perform computations over natural language representations to solve tasks that require symbolic and numeric reasoning? We propose natural language embedded programs (NLEP) as a unifying framework for addressing math/symbolic reasoning, natural language understanding, and instruction following tasks. Our approach prompts a language model to generate full Python programs that define functions over data structures which contain natural language representations of structured knowledge. A Python interpreter then executes the generated code and prints the output. Despite using a task-general prompt, we find that this approach can improve upon strong baselines across a range of different tasks including math and symbolic reasoning, text classification, question answering, and instruction following. We found that the generated programs are interpretable since they outline the exact reasoning process followed by the program interpreter.
Tianhua Zhang, Jiaxin Ge, Hongyin Luo, Yung-Sung Chuang, Mingye Gao, Yuan Gong, Xixin Wu, Yoon Kim, Helen Meng, James Glass
arXiv:2309.10814 · cs.CL · submitted Sep 19, 2023 · updated Mar 29, 2024
abstract · pdf · html · NAACL 2024
Paper: https://arxiv.org/abs/2309.10814 An automatic NLEP generation toolkit is opensourced: https://github.com/luohongyin/langcode
Example Colab notebook is included in the Github repo.
This work introduces the following features of NLEP
1. NLEP is a full python program that prints the target response of LLMs. 2. Task-general NLEP prompting outperforms task-specific chain-of-thought prompting on math, symbolic, and natural language. 3. Enable the chain-of-thought reasoning ability of small models (RoBERTa) on text classification 4. Hierarchical instructing via program completion.