about
The Impact of Prompt Programming on Function-Level Code Generation (arxiv.org)
3 points by bobrenjc93 on Feb 10, 2025 | hide | past | pdf | discuss on HN

In plain words: Tested five tricks—examples, a persona, step-by-step reasoning, function signature, and allowed packages—on three code AIs with 7,072 prompts to see which give correct code. Some tricks changed the code, but stacking several did not reliably help, and more correctness meant lower quality.

Abstract

Large Language Models (LLMs) are increasingly used by software engineers for code generation. However, limitations of LLMs such as irrelevant or incorrect code have highlighted the need for prompt programming (or prompt engineering) where engineers apply specific prompt techniques (e.g., chain-of-thought or input-output examples) to improve the generated code. While some prompt techniques have been studied, the impact of different techniques -- and their interactions -- on code generation is still not fully understood. In this study, we introduce CodePromptEval, a dataset of 7072 prompts designed to evaluate five prompt techniques (few-shot, persona, chain-of-thought, function signature, list of packages) and their effect on the correctness, similarity, and quality of complete functions generated by three LLMs (GPT-4o, Llama3, and Mistral). Our findings show that while certain prompt techniques significantly influence the generated code, combining multiple techniques does not necessarily improve the outcome. Additionally, we observed a trade-off between correctness and quality when using prompt techniques. Our dataset and replication package enable future research on improving LLM-generated code and evaluating new prompt techniques.

Ranim Khojah, Francisco Gomes de Oliveira Neto, Mazen Mohamad, Philipp Leitner
arXiv:2412.20545 · cs.SE, cs.CL, cs.HC, cs.LG · submitted Dec 29, 2024 · updated Jul 8, 2025
abstract · pdf · html · Accepted at Transactions on Software Engineering (TSE). CodePromptEval dataset and replication package on GitHub: https://github.com/icetlab/CodePromptEval

add comment on HN