about
PanGu-Coder2: SOTA for Code LLMs on 15B (arxiv.org)
8 points by fgfm on Jul 28, 2023 | hide | past | pdf | 2 comments on HN

In plain words: Instead of training a coding model on human-written answers, this system ranks the model's candidate programs using tests and a teacher's feedback, then trains it on the highest-ranked ones. It solved 62.2% of standard coding problems on the first try, beating earlier code models.

Abstract · PanGu-Coder2: Boosting Large Language Models for Code with Ranking Feedback

Large Language Models for Code (Code LLM) are flourishing. New and powerful models are released on a weekly basis, demonstrating remarkable performance on the code generation task. Various approaches have been proposed to boost the code generation performance of pre-trained Code LLMs, such as supervised fine-tuning, instruction tuning, reinforcement learning, etc. In this paper, we propose a novel RRTF (Rank Responses to align Test&Teacher Feedback) framework, which can effectively and efficiently boost pre-trained large language models for code generation. Under this framework, we present PanGu-Coder2, which achieves 62.20% pass@1 on the OpenAI HumanEval benchmark. Furthermore, through an extensive evaluation on CoderEval and LeetCode benchmarks, we show that PanGu-Coder2 consistently outperforms all previous Code LLMs.

Bo Shen, Jiaxin Zhang, Taihong Chen, Daoguang Zan, Bing Geng, An Fu, Muhan Zeng, Ailun Yu, Jichuan Ji, Jingyang Zhao, Yuenan Guo, Qianxiang Wang
arXiv:2307.14936 · cs.CL, cs.AI, cs.LG, cs.PL, cs.SE · submitted Jul 27, 2023
abstract · pdf · html · Preprint

add comment on HN

There is a new code LLM in town, and with only 15B params, it reaches 62.20% pass@1 on HumanEval!
Boom