about
Detecting Language Model Attacks with Perplexity (arxiv.org)
1 point by diogotozzi on Aug 29, 2023 | hide | past | pdf | discuss on HN

In plain words: Attackers append gibberish to a prompt to trick a chatbot into giving dangerous instructions. Scoring how confused a small language model is by the wording flags these, but a simple classifier using that score plus prompt length caught most attacks without scoring's false alarms.

Abstract

A novel hack involving Large Language Models (LLMs) has emerged, exploiting adversarial suffixes to deceive models into generating perilous responses. Such jailbreaks can trick LLMs into providing intricate instructions to a malicious user for creating explosives, orchestrating a bank heist, or facilitating the creation of offensive content. By evaluating the perplexity of queries with adversarial suffixes using an open-source LLM (GPT-2), we found that they have exceedingly high perplexity values. As we explored a broad range of regular (non-adversarial) prompt varieties, we concluded that false positives are a significant challenge for plain perplexity filtering. A Light-GBM trained on perplexity and token length resolved the false positives and correctly detected most adversarial attacks in the test set.

Gabriel Alon, Michael Kamfonas
arXiv:2308.14132 · cs.CL, cs.AI, cs.CR, cs.LG · submitted Aug 27, 2023 · updated Nov 7, 2023
abstract · pdf · html

add comment on HN