about
How to Hack Transformers: Steering LLMs via Prompts, States, and Weight Edits (arxiv.org)
2 points by WASDAai on Sep 8, 2025 | hide | past | pdf | 1 comment on HN

In plain words: A framework steers language models by editing the prompt, its internal signals, or its weights, seeking the smallest change that gives the wanted behavior. It flipped sentiment and fixed facts over 90% of the time without hurting performance, though broader edits bring side effects.

Abstract · Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions

Transformer-based language models excel in NLP tasks, but fine-grained control remains challenging. This paper explores methods for manipulating transformer models through principled interventions at three levels: prompts, activations, and weights. We formalize controllable text generation as an optimization problem addressable via prompt engineering, parameter-efficient fine-tuning, model editing, and reinforcement learning. We introduce a unified framework encompassing prompt-level steering, activation interventions, and weight-space edits. We analyze robustness and safety implications, including adversarial attacks and alignment mitigations. Theoretically, we show minimal weight updates can achieve targeted behavior changes with limited side-effects. Empirically, we demonstrate >90% success in sentiment control and factual edits while preserving base performance, though generalization-specificity trade-offs exist. We discuss ethical dual-use risks and the need for rigorous evaluation. This work lays groundwork for designing controllable and robust language models.

Faruk Alpay, Taylan Alpay
arXiv:2509.04549 · cs.CL, cs.AI · submitted Sep 4, 2025
abstract · pdf · html · 13 pages

add comment on HN

TL;DR: The paper shows how you can steer LLMs by messing with prompts, hidden states, or weight edits—and warns that the same tricks can be used maliciously.