about
Demystifying GPT Self-Repair for Code Generation (arxiv.org)
2 points by famouswaffles on Jul 4, 2023 | hide | past | pdf | 1 comment on HN

In plain words: They tested whether letting a model debug its own code helps on coding problems, counting the cost of extra attempts. Gains were often small or missing, but giving the model better feedback, from a stronger model or humans, improved results much more.

Abstract · Is Self-Repair a Silver Bullet for Code Generation?

Large language models have shown remarkable aptitude in code generation, but still struggle to perform complex tasks. Self-repair -- in which the model debugs and repairs its own code -- has recently become a popular way to boost performance in these settings. However, despite its increasing popularity, existing studies of self-repair have been limited in scope; in many settings, its efficacy thus remains poorly understood. In this paper, we analyze Code Llama, GPT-3.5 and GPT-4's ability to perform self-repair on problems taken from HumanEval and APPS. We find that when the cost of carrying out repair is taken into account, performance gains are often modest, vary a lot between subsets of the data, and are sometimes not present at all. We hypothesize that this is because self-repair is bottlenecked by the model's ability to provide feedback on its own code; using a stronger model to artificially boost the quality of the feedback, we observe substantially larger performance gains. Similarly, a small-scale study in which we provide GPT-4 with feedback from human participants suggests that even for the strongest models, self-repair still lags far behind what can be achieved with human-level debugging.

Theo X. Olausson, Jeevana Priya Inala, Chenglong Wang, Jianfeng Gao, Armando Solar-Lezama
arXiv:2306.09896 · cs.CL, cs.AI, cs.PL, cs.SE · submitted Jun 16, 2023 · updated Feb 2, 2024
abstract · pdf · html · Accepted to ICLR 2024. Added additional Code Llama experiments and fixed a data processing error harming Code Llama's reported self-repair performance on HumanEval

add comment on HN

What I thought particularly interesting

"....With this evaluation strategy, we find that the effectiveness of self-repair is only seen in GPT-4. We also observe that self-repair is bottlenecked by the feedback stage; using GPT-4 to give feedback on the programs generated by GPT-3.5 and using expert human programmers to give feedback on the programs generated by GPT-4, we unlock significant performance gains."