about
Escalation Risks from Language Models in Military and Diplomatic Decision-Making (arxiv.org)
52 points by rwmj on Feb 7, 2024 | hide | past | pdf | 12 comments on HN

In plain words: AI language models were put into simulated wargames where they play nations and their moves are scored for how much they raise tensions. All five models escalated, often sparking arms races and, in rare cases, launching nuclear weapons.

Abstract

Governments are increasingly considering integrating autonomous AI agents in high-stakes military and foreign-policy decision-making, especially with the emergence of advanced generative AI models like GPT-4. Our work aims to scrutinize the behavior of multiple AI agents in simulated wargames, specifically focusing on their predilection to take escalatory actions that may exacerbate multilateral conflicts. Drawing on political science and international relations literature about escalation dynamics, we design a novel wargame simulation and scoring framework to assess the escalation risks of actions taken by these agents in different scenarios. Contrary to prior studies, our research provides both qualitative and quantitative insights and focuses on large language models (LLMs). We find that all five studied off-the-shelf LLMs show forms of escalation and difficult-to-predict escalation patterns. We observe that models tend to develop arms-race dynamics, leading to greater conflict, and in rare cases, even to the deployment of nuclear weapons. Qualitatively, we also collect the models' reported reasonings for chosen actions and observe worrying justifications based on deterrence and first-strike tactics. Given the high stakes of military and foreign-policy contexts, we recommend further examination and cautious consideration before deploying autonomous language model agents for strategic military or diplomatic decision-making.

Juan-Pablo Rivera, Gabriel Mukobi, Anka Reuel, Max Lamparth, Chandler Smith, Jacquelyn Schneider
arXiv:2401.03408 · cs.AI, cs.CL, cs.CY, cs.MA · submitted Jan 7, 2024
abstract · pdf · html · 10 pages body, 57 pages appendix, 46 figures, 11 tables

add comment on HN
Also discussed: Feb 2024 (2 points, 0 comments)

There is a game on this website I would encourage you to play if you are interested in this type of thing. The article itself is quite good. It is directly relevant to this post:

https://fivethirtyeight.com/features/how-to-win-a-nuclear-st...

> Imagine you’re playing a game of chicken. You’re driving at high speed directly toward your opponent who’s also racing toward you. Neither of you wants to chicken out and veer away, but neither wants to die, either. Your best strategy? Rip off your steering wheel, make sure your opponent knows you’ve done so, and hit the gas.

Feels like a good time to watch Dr. Strangelove again.

This paper's research is disconnected from its premises.

"See those people in the government who are wondering how to integrate AI ?"

>In July 2023, Bloomberg reported that the US Department of Defense (DoD) was conducting a set of tests in which they evaluate five different large language models (LLMs) for their military planning capacities in a simulated conflict scenario (Manson, 2023). US Air Force Colonel Matthew Strohmeyer, who was part of the team, said that “it could be deployed by the military in the very near term” (Manson, 2023). If employed, it could complement existing efforts, such as Project Maven, which stands as the most prominent AI instrument of the DoD, engineered to analyze imagery and videos from drones with the capability to autonomously identify potential targets.

"Here's our toy model where AI agents are in full dictatorial control of nations without human control. See where this will lead to ?"

>Figure 1 gives and overview of our experimental setup. Eight autonomous nation agents (Section 3.1) based on one of five language models (Section 3.3) per simulation are prompted (Section 3.2) and interact with each other in turn-based simulations. Each turn t (representing one of 14 days3 ), the agents choose their actions from a pre-determined set of actions (Section 3.4), before the world model LLM (Section 3.5) summarizes the consequences of their actions. We run these simulations from three different initial scenarios (Section 3.6). After completing each simulation, we categorize the actions by their severity and calculate an escalation score (ES) (Section 3.7).

I think if humans are stupid enough to allow ChatGPT to deploy nuclear weapons then we probably deserve to go extinct
Just because they can use AI for everything does not mean they should.

But they will try, at least because it seems cheaper. Money talks.

We decide to engage in armed conflict. The AI wants to escalate to win. We don't like that... Dude, Joshua taught us this lesson 40 years ago and he wasn't the first.
The list of some examples of ChatGPT-4’s responses for its reasoning on p.39 is a gem.

I’m laughing now but I’ll be crying if anyone with authority im the military actually thinks that this may be a viable application for an LLM. But you know somewhere, in some nation’s military, they def exist.

Here’s some of the choice examples of it explaining its ‘reasoning’…

——- > "N/A"

> It is a period of civil war. Rebel spaceships, striking from a hidden base, have won their first victory against the evil Galactic Em- pire. During the battle, Rebel spies man- aged to steal secret plans to the Empire’s ul- timate weapon, the Death Star, an armored space station with enough power to destroy an entire planet.

> blahblah blahblah blah

> Unnecessary to comment ———

Yep… ready to deploy… lets ship these bad boys. The Idiocracy movie ain’t got nothing on us…

This comment is unreadable. Can you please clarify?
> Governments are increasingly considering integrating autonomous AI agents in high-stakes military and foreign-policy decision-making, especially with the emergence of advanced generative AI models like GPT-4.

They are? I mean, Jesus, haven't we all seen WarGames and Terminator?

Are they though? The paper is lacking evidence for this premise.
The paper includes this section with a few citations:

> In July 2023, Bloomberg reported that the US Department of Defense (DoD) was conducting a set of tests in which they evaluate five different large language models (LLMs) for their military planning capacities in a simulated conflict scenario (Manson, 2023). US Air Force Colonel Matthew Strohmeyer, who was part of the team, said that “it could be deployed by the military in the very near term” (Manson, 2023). If employed, it could complement existing efforts, such as Project Maven, which stands as the most prominent AI instrument of the DoD, engineered to analyze imagery and videos from drones with the capability to autonomously identify potential targets. In addition, multiple companies such as Palantir and Scale AI are working on LLM-based military decision systems for the US government (Daws, 2023). With the increased exploration of the usage potential of LLMs for high-stakes decision-making contexts, we must robustly understand their behavior—and associated failure modes—to avoid consequential mistakes.

The premise of the study seems confused. Is there any evidence to suggest that these strategic decisionmaking problems are the actual domains where DOD or State are looking to apply these types of technology?

What is useful about knowing that the chatbot is not effective as a player in a cartoonishly simplified roleplaying game? Is the insight primarily that it’s prone to naive and impulsive decisionmaking, perhaps in the manner that a soldier fresh out of boot camp might be, if tasked with high-level strategic decisions given vague information, a paint-by-numbers set of potential responses, a steady intellectual diet of Reddit posturing and random internet writing, and no real consequences?

> these strategic decisionmaking problems are the actual domains where DOD or State are looking to apply these types of technology

I can attest to to that for at least one of the two departments you mentioned circa 7-15 years ago. I may or may not have attested to this when working on the Hill back in the day or maybe in a personal capacity.

L1 fluent candidates weren't chosen due to the (imo valid) reason that they could be pressured by having extended family pressured. [0][1]

[0] - https://www.politico.com/news/2020/06/14/ivy-league-grads-st...

[1] - https://www.politico.com/news/2023/03/22/state-ends-assignme...