about
Will GPT-4 Run Doom? (arxiv.org)
5 points by Hard_Space on Mar 11, 2024 | hide | past | pdf | 2 comments on HN

In plain words: GPT-4 plays Doom by writing a description of each screenshot and choosing moves from that text, with a few instructions. It opened doors, fought enemies, and navigated well enough to play, though AI trained on games plays better — and GPT-4 needed no training.

Abstract · Will GPT-4 Run DOOM?

We show that GPT-4's reasoning and planning capabilities extend to the 1993 first-person shooter Doom. This large language model (LLM) is able to run and play the game with only a few instructions, plus a textual description--generated by the model itself from screenshots--about the state of the game being observed. We find that GPT-4 can play the game to a passable degree: it is able to manipulate doors, combat enemies, and perform pathing. More complex prompting strategies involving multiple model calls provide better results. While further work is required to enable the LLM to play the game as well as its classical, reinforcement learning-based counterparts, we note that GPT-4 required no training, leaning instead on its own reasoning and observational capabilities. We hope our work pushes the boundaries on intelligent, LLM-based agents in video games. We conclude by discussing the ethical implications of our work.

Adrian de Wynter
arXiv:2403.05468 · cs.CL, cs.AI, cs.CV · submitted Mar 8, 2024
abstract · pdf · html

add comment on HN
Also discussed: Mar 2024 (5 points, 0 comments)

This is a nice read and an interesting concept.

I'm also sure that improving an LLM with some extra sauce that doesn't rely on token generation and making it better at handling such tasks as playing Doom is the next step.

With LLMs we've built a better "language" brain than a human, now we need to build and research all the rest.

Just pumping more money in LLMs will yield low returns, but it's the hype of the moment, can't wait until it crashes so we are back to real science.

I bet Sora can "run" Doom.