about
Can You Trust Code Copilots? Evaluating LLMs from a Code Security Perspec (arxiv.org)
11 points by badmonster on May 18, 2025 | hide | past | pdf | 2 comments on HN

In plain words: They built a test suite covering code writing, vulnerability spotting, naming and fixing, plus an automated reviewer that flags flaws like human experts. Across 20 coding assistants, models spotted vulnerable code well but wrote insecure code and rarely fixed or named flaws.

Abstract · Can You Really Trust Code Copilots? Evaluating Large Language Models from a Code Security Perspective

Code security and usability are both essential for various coding assistant applications driven by large language models (LLMs). Current code security benchmarks focus solely on single evaluation task and paradigm, such as code completion and generation, lacking comprehensive assessment across dimensions like secure code generation, vulnerability repair and discrimination. In this paper, we first propose CoV-Eval, a multi-task benchmark covering various tasks such as code completion, vulnerability repair, vulnerability detection and classification, for comprehensive evaluation of LLM code security. Besides, we developed VC-Judge, an improved judgment model that aligns closely with human experts and can review LLM-generated programs for vulnerabilities in a more efficient and reliable way. We conduct a comprehensive evaluation of 20 proprietary and open-source LLMs. Overall, while most LLMs identify vulnerable codes well, they still tend to generate insecure codes and struggle with recognizing specific vulnerability types and performing repairs. Extensive experiments and qualitative analyses reveal key challenges and optimization directions, offering insights for future research in LLM code security.

Yutao Mou, Xiao Deng, Yuxiao Luo, Shikun Zhang, Wei Ye
arXiv:2505.10494 · cs.CL · submitted May 15, 2025
abstract · pdf · html · Accepted by ACL2025 Main Conference

add comment on HN

Any headline containing a question can be answered in the negative.

My colleague sent me an awesome suggestion by copilot. It introduced a classic SQL injection vulnerability. I copy pasted it into Claude and said “any problems with this?” Yep! Critical sql injection vulnerability!

So even when they know they can’t be trusted (granted this is two different models). But the secret is that they often do know best practices and they simply don’t follow them. I personally have experienced half a dozen pretty severe problems with ai code.

They’re like lying, idiot-savant teenagers. They will 100% forget to propagate variable names (and claim to have checked for that when challenged), fail to follow best security practices, lie, claim to have done things when they have not, fail to think carefully.

And they can write 2,000 lines of reasonably good code an hour in any language you like. And they’re not bad as pair programmers.

Oh and they can review your code reasonably well.

Maybe with some good code security cursor rules they can get better, but I’ll believe it when I see it.

So no, don’t trust them.

answer: no