In plain words: Build powerful AI so it is proven to follow rules humans write, using AI to check the code and inspect its inner workings. This approach is argued to be the only one that guarantees safe, controlled AGI, and to become technically possible soon.
Abstract
We describe a path to humanity safely thriving with powerful Artificial General Intelligences (AGIs) by building them to provably satisfy human-specified requirements. We argue that this will soon be technically feasible using advanced AI for formal verification and mechanistic interpretability. We further argue that it is the only path which guarantees safe controlled AGI. We end with a list of challenge problems whose solution would contribute to this positive outcome and invite readers to join in this work.
Max Tegmark, Steve Omohundro
arXiv:2309.01933 · cs.CY, cs.AI, cs.LG · submitted Sep 5, 2023
abstract · pdf · html · 17 pages
Well, if this is correct, we should routinely be proving safety of ordinary systems long before we get to AI. I'd like to see formally verified routers, firewalls, mailers, and DNS servers, all of which have definable correct behavior, in wide use. That's probably possible now.
Defining safe behavior for a LLM is a much harder problem. The paper handwaves this.
* Mortal AI. Has death date. Does not require proof, just a hardware timer or limited battery life.
* Geofenced AI. Only useful for mobile machines. Not helpful against things which can communicate.
* Throttled AI. You have to keep putting in crypto tokens to keep it going. OK, whatever.
* AI kill switch. Off switch.
* Asimov-style laws. Not an inherently bad idea, but way too ambiguous to rigorously formalize. Go read Asimov's robot books again. Useful metric: what similar set of bright-line constraints could usefully be enforced on corporations?
It's worth bearing in mind that most of the problems of regulating AIs apply to regulating corporations, which can be thought of as AIs with slow internal data transfer.