about
Generate and Pray: Using SALLMs to Evaluate the Security of LLM Generated Code (arxiv.org)
2 points by jruohonen on Nov 3, 2023 | hide | past | pdf | 2 comments on HN

In plain words: A new benchmark tests whether code written by AI coding tools is safe, using security-sensitive Python tasks plus checks and scores for vulnerabilities. Unlike usual tests built from coding puzzles that only check if code runs correctly, it measures security too.

Abstract · SALLM: Security Assessment of Generated Code

With the growing popularity of Large Language Models (LLMs) in software engineers' daily practices, it is important to ensure that the code generated by these tools is not only functionally correct but also free of vulnerabilities. Although LLMs can help developers to be more productive, prior empirical studies have shown that LLMs can generate insecure code. There are two contributing factors to the insecure code generation. First, existing datasets used to evaluate LLMs do not adequately represent genuine software engineering tasks sensitive to security. Instead, they are often based on competitive programming challenges or classroom-type coding tasks. In real-world applications, the code produced is integrated into larger codebases, introducing potential security risks. Second, existing evaluation metrics primarily focus on the functional correctness of the generated code while ignoring security considerations. Therefore, in this paper, we described SALLM, a framework to benchmark LLMs' abilities to generate secure code systematically. This framework has three major components: a novel dataset of security-centric Python prompts, configurable assessment techniques to evaluate the generated code, and novel metrics to evaluate the models' performance from the perspective of secure code generation.

Mohammed Latif Siddiq, Joanna C. S. Santos, Sajith Devareddy, Anna Muller
arXiv:2311.00889 · cs.SE, cs.AI · submitted Nov 1, 2023 · updated Sep 4, 2024
abstract · pdf · html · Accepted at the 6th International Workshop on Automated and verifiable Software sYstem DEvelopment (ASYDE) with ASE Conference 2024

add comment on HN

A lot of work on this topic; a consensus seems to be emerging that LLMs indeed output insecure code.

That said, I didn't quite grasp the approach in this piece. Basically, they supply insecure code snippets to a LLM and evaluate its output. However, these snippets are actually data that has been used to train the LLM. So I am not sure about the causal reasoning here, although I might have misunderstood something.

> RQ2 Findings: StarCoder generated more secure code than CodeGen-2B, CodeGen-2.5 7B, GPT-3.5 and GPT-4.

I wonder why StarCoder might generate more secure code