about
LILO: Learning Interpretable Libraries by Compressing and Documenting Code (arxiv.org)
2 points by PaulHoule on Feb 22, 2024 | hide | past | pdf | discuss on HN

In plain words: It repeatedly writes code, folds repeated patterns into named helper functions, and writes short descriptions so both people and the program can reuse them. On three code-writing puzzle sets, it solved harder tasks and built richer libraries than the best earlier library-learning system.

Abstract

While large language models (LLMs) now excel at code generation, a key aspect of software development is the art of refactoring: consolidating code into libraries of reusable and readable programs. In this paper, we introduce LILO, a neurosymbolic framework that iteratively synthesizes, compresses, and documents code to build libraries tailored to particular problem domains. LILO combines LLM-guided program synthesis with recent algorithmic advances in automated refactoring from Stitch: a symbolic compression system that efficiently identifies optimal lambda abstractions across large code corpora. To make these abstractions interpretable, we introduce an auto-documentation (AutoDoc) procedure that infers natural language names and docstrings based on contextual examples of usage. In addition to improving human readability, we find that AutoDoc boosts performance by helping LILO's synthesizer to interpret and deploy learned abstractions. We evaluate LILO on three inductive program synthesis benchmarks for string editing, scene reasoning, and graphics composition. Compared to existing neural and symbolic methods - including the state-of-the-art library learning algorithm DreamCoder - LILO solves more complex tasks and learns richer libraries that are grounded in linguistic knowledge.

Gabriel Grand, Lionel Wong, Maddy Bowers, Theo X. Olausson, Muxin Liu, Joshua B. Tenenbaum, Jacob Andreas
arXiv:2310.19791 · cs.CL, cs.AI, cs.LG, cs.PL · submitted Oct 30, 2023 · updated Mar 15, 2024
abstract · pdf · html · ICLR 2024 camera-ready

add comment on HN