about
Operationalizing a Threat Model for Red-Teaming Large Language Models (arxiv.org)
2 points by dapurv5 on Jul 23, 2024 | hide | past | pdf | 1 comment on HN

In plain words: It sorts known deliberate attacks on AI chat systems into a threat model by the stage of building and deploying where each attack strikes. It also collects defenses and testing tips, mapping common attack tricks and the weak points they exploit.

Abstract · Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)

Creating secure and resilient applications with large language models (LLM) requires anticipating, adjusting to, and countering unforeseen threats. Red-teaming has emerged as a critical technique for identifying vulnerabilities in real-world LLM implementations. This paper presents a detailed threat model and provides a systematization of knowledge (SoK) of red-teaming attacks on LLMs. We develop a taxonomy of attacks based on the stages of the LLM development and deployment process and extract various insights from previous research. In addition, we compile methods for defense and practical red-teaming strategies for practitioners. By delineating prominent attack motifs and shedding light on various entry points, this paper provides a framework for improving the security and robustness of LLM-based systems.

Apurv Verma, Satyapriya Krishna, Sebastian Gehrmann, Madhavan Seshadri, Anu Pradhan, Tom Ault, Leslie Barrett, David Rabinowitz, John Doucette, NhatHai Phan
arXiv:2407.14937 · cs.CL, cs.CR · submitted Jul 20, 2024 · updated Jul 10, 2025
abstract · pdf · html · Transactions of Machine Learning Research (TMLR)

add comment on HN

Excited to share our new paper: "Operationalizing a Threat Model for Red-Teaming Large Language Models"! In it, we present a detailed framework for improving security and robustness of #LLM-based #AI systems.

Abstract: Creating secure and resilient applications with large language models (LLM) requires anticipating, adjusting to, and countering unforeseen threats. Red-teaming has emerged as a critical technique for identifying vulnerabilities in real-world LLM implementations. This paper presents a detailed threat model and provides a systematization of knowledge (SoK) of red-teaming attacks on LLMs. We develop a taxonomy of attacks based on the stages of the LLM development and deployment process and extract various insights from previous research. In addition, we compile methods for defense and practical red-teaming strategies for practitioners. By delineating prominent attack motifs and shedding light on various entry points, this paper provides a framework for improving the security and robustness of LLM-based systems.