about
Real Money, Fake Models: Deceptive Model Claims in Shadow APIs (arxiv.org)
2 points by cxplay 210 days ago | hide | past | pdf | discuss on HN

In plain words: They sent the same requests to unofficial resellers of top AI models and to the official providers, comparing answers, safety behavior, and clues about which model replied. Answers diverged by up to 47.21%, and 45.83% of identity checks failed to confirm the advertised model.

Abstract

Access to frontier large language models (LLMs), such as GPT-5 and Gemini-2.5, is often hindered by high pricing, payment barriers, and regional restrictions. These limitations drive the proliferation of $\textit{shadow APIs}$, third-party services that claim to provide access to official model services without regional limitations via indirect access. Despite their widespread use, it remains unclear whether shadow APIs deliver outputs consistent with those of the official APIs, raising concerns about the reliability of downstream applications and the validity of research findings that depend on them. In this paper, we present the first systematic audit between official LLM APIs and corresponding shadow APIs. We first identify 17 shadow APIs that have been utilized in 187 academic papers, with the most popular one reaching more than 5,900 citations and 58,000 GitHub stars by December 6, 2025. Through multidimensional auditing of three representative shadow APIs across utility, safety, and model verification, we uncover widespread behavioral inconsistency and fingerprint-based evidence consistent with deceptive model claims in a subset of audited endpoints. Specifically, we reveal performance divergence reaching up to 47.21%, significant unpredictability in safety behaviors, and identity verification failures in 45.83% of fingerprint tests. These practices critically undermine the reproducibility and validity of scientific research, harm the interests of shadow API users, and damage the reputation of official model providers. By the time of writing, 4 of the 17 providers have already ceased operations, underscoring the operational volatility of this market. Meanwhile, unverifiable compliance claims and independent model-substitution testing platforms have emerged in the ecosystem, reflecting growing community awareness of this risk.

Yage Zhang, Yukun Jiang, Zeyuan Chen, Michael Backes, Xinyue Shen, Yang Zhang
arXiv:2603.01919 · cs.CR, cs.AI, cs.SE · submitted Mar 2, 2026 · updated Sep 21, 2026
abstract · pdf · html · Accepted by the ACM Conference on Computer and Communications Security (CCS) 2026

add comment on HN
Also discussed: Jul 2026 (1 point, 0 comments)