about
Security Degradation in Iterative AI Code Generation (arxiv.org)
1 point by chillax on Oct 3, 2025 | hide | past | pdf | discuss on HN

In plain words: Tested 400 code samples through 40 rounds of AI-driven "improvements" using four different ways of asking for changes, tracking how security flaws evolved. Critical vulnerabilities rose 37.6% after just five rounds, showing repeated AI polishing can make code less safe than leaving it alone.

Abstract · Security Degradation in Iterative AI Code Generation -- A Systematic Analysis of the Paradox

The rapid adoption of Large Language Models(LLMs) for code generation has transformed software development, yet little attention has been given to how security vulnerabilities evolve through iterative LLM feedback. This paper analyzes security degradation in AI-generated code through a controlled experiment with 400 code samples across 40 rounds of "improvements" using four distinct prompting strategies. Our findings show a 37.6% increase in critical vulnerabilities after just five iterations, with distinct vulnerability patterns emerging across different prompting approaches. This evidence challenges the assumption that iterative LLM refinement improves code security and highlights the essential role of human expertise in the loop. We propose practical guidelines for developers to mitigate these risks, emphasizing the need for robust human validation between LLM iterations to prevent the paradoxical introduction of new security issues during supposedly beneficial code "improvements".

Shivani Shukla, Himanshu Joshi, Romilla Syed
arXiv:2506.11022 · cs.SE, cs.AI, cs.CL, cs.CR, cs.LG · submitted May 19, 2025 · updated Sep 26, 2025
abstract · pdf · html · Keywords - Large Language Models, Security Vulnerabilities, AI-Generated Code, Iterative Feedback, Software Security, Secure Coding Practices, Feedback Loops, LLM Prompting Strategies

add comment on HN