about
A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs (arxiv.org)
31 points by rntn on Oct 1, 2024 | hide | past | pdf | 9 comments on HN

In plain words: Code-writing AI models sometimes invent package names that don't exist, letting an attacker who registers them sneak malicious code into projects. In hundreds of thousands of code samples, open-source models invented fake packages 21.7% of the time, more than commercial ones; simple fixes cut it sharply.

Abstract · We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs

The reliance of popular programming languages such as Python and JavaScript on centralized package repositories and open-source software, combined with the emergence of code-generating Large Language Models (LLMs), has created a new type of threat to the software supply chain: package hallucinations. These hallucinations, which arise from fact-conflicting errors when generating code using LLMs, represent a novel form of package confusion attack that poses a critical threat to the integrity of the software supply chain. This paper conducts a rigorous and comprehensive evaluation of package hallucinations across different programming languages, settings, and parameters, exploring how a diverse set of models and configurations affect the likelihood of generating erroneous package recommendations and identifying the root causes of this phenomenon. Using 16 popular LLMs for code generation and two unique prompt datasets, we generate 576,000 code samples in two programming languages that we analyze for package hallucinations. Our findings reveal that that the average percentage of hallucinated packages is at least 5.2% for commercial models and 21.7% for open-source models, including a staggering 205,474 unique examples of hallucinated package names, further underscoring the severity and pervasiveness of this threat. To overcome this problem, we implement several hallucination mitigation strategies and show that they are able to significantly reduce the number of package hallucinations while maintaining code quality. Our experiments and findings highlight package hallucinations as a persistent and systemic phenomenon while using state-of-the-art LLMs for code generation, and a significant challenge which deserves the research community's urgent attention.

Joseph Spracklen, Raveen Wijewickrama, A H M Nazmus Sakib, Anindya Maiti, Bimal Viswanath, Murtuza Jadliwala
arXiv:2406.10279 · cs.SE, cs.AI, cs.CR, cs.LG · submitted Jun 12, 2024 · updated Mar 2, 2025
abstract · pdf · html · To appear in the 2025 USENIX Security Symposium. 22 pages, 14 figures, 8 tables. Edited from original version for submission to a different conference. No change to original results or findings

add comment on HN

This seems to be a problem I frequently run into. The LLM tends to suggest using libraries that don't seem to ever have existed as far as I can find.
I always specify the libraries I want it to use. If I'm not sure, or it introduces something I'm not familiar with, I spend a few minutes comparing similar libraries so I can.
Just ask the program to implement those packages too? If the packages don't exist, maybe they should exist.
I hope you're joking
> One course of action that we chose not to pursue for ethical reasons was publishing actual packages using hallucinated package names to PyPI

I mean, this makes sense from a security perspective. But from a language usage perspective, if there is a missing package that would be super-useful, then implementing and publishing that package would be a win.

I'm curious what the package names were, they seem to have deliberately omitted any package names. Maybe there are some good package ideas in the 19% of names that were hallucinated by multiple models.

No please don’t spam the repo with ai trash aliases!!!

The correct and idiomatic way to implement this is to redirect the insane ai guessing to a local proxy which can perform the required search and replace, if that truly is all you need (it is not)

By mistakenly declaring the existence of certain packages at scale, the model causes those packages to be created and published. What initially seemed like a hallucination was in fact hyperstition...
> hyperstition

Not to be confused with substition.

> 71-hour Ahmed was not superstitious. He was substitious, which put him in a minority among humans. He didn't believe in the things everyone believed in but which nevertheless weren't true. He believed instead in the things that were true in which no one else believed. There are many such substitions, ranging from ‘It’ll get better if you don’t pick at it’ all the way up to ‘Sometimes things just happen.’

-- Jingo by Terry Pratchett

It would be useful for Software Composition Analysis tools to know the list so they can be flagged. But that can happen on a direct basis instead of publishing the list.