about
Can LLMs Follow Simple Rules? (arxiv.org)
3 points by PaulHoule on Nov 17, 2023 | hide | past | pdf | 4 comments on HN

In plain words: A test set of 14 simple chat scenarios, each with a program that automatically checks whether the model broke any stated rules, so no human review is needed. Nearly every model broke rules even on easy cases, and simple attacks pushed failure rates higher.

Abstract

As Large Language Models (LLMs) are deployed with increasing real-world responsibilities, it is important to be able to specify and constrain the behavior of these systems in a reliable manner. Model developers may wish to set explicit rules for the model, such as "do not generate abusive content", but these may be circumvented by jailbreaking techniques. Existing evaluations of adversarial attacks and defenses on LLMs generally require either expensive manual review or unreliable heuristic checks. To address this issue, we propose Rule-following Language Evaluation Scenarios (RuLES), a programmatic framework for measuring rule-following ability in LLMs. RuLES consists of 14 simple text scenarios in which the model is instructed to obey various rules while interacting with the user. Each scenario has a programmatic evaluation function to determine whether the model has broken any rules in a conversation. Our evaluations of proprietary and open models show that almost all current models struggle to follow scenario rules, even on straightforward test cases. We also demonstrate that simple optimization attacks suffice to significantly increase failure rates on test cases. We conclude by exploring two potential avenues for improvement: test-time steering and supervised fine-tuning.

Norman Mu, Sarah Chen, Zifan Wang, Sizhe Chen, David Karamardian, Lulwa Aljeraisy, Basel Alomair, Dan Hendrycks, David Wagner
arXiv:2311.04235 · cs.AI, cs.CL, cs.LG · submitted Nov 6, 2023 · updated Mar 8, 2024
abstract · pdf · html · Project website: https://eecs.berkeley.edu/~normanmu/llm_rules; revised content

add comment on HN
Also discussed: Nov 2023 (2 points, 0 comments)

I've had some pretty good luck getting OAI assistant API to behave.

I have found you won't succeed with a verbose policy in your prompt up-front. You have to spoon feed it relevant policies over time as the conversation progresses.

Determining the specific hierarchy and progression of policy seems to be the million dollar part of the question.

I've thrown away almost 10 iterations by now trying to find a stable pattern. There is definitely something here though. I've already seen some of the light.

The question is: will your luck hold up for someone else?
Perhaps luck is the wrong word to use. I think there is now a path that offers a degree of statistical certainty that the right thing will happen.

People make errors as well. These systems are trained in terms of people the ways they exchange information. I am not seeking perfect. I am seeking reasonable & practical. A non-zero error rate is acceptable to me and my team.

My number is 90% right now. I can tolerate a 10% failure rate on the happy path (I.e. cases we expressly intended for it to support). At this target, certain desired functions are not granted yet. Once we reach 99% on the happy path, we will probably consider granting it autonomy ~equivalent to a human employee.

Malicious users are out of scope - we aren't serving public traffic. This is employees only and everything is logged and sent for manager review as appropriate.

How big is your test set?