In plain words: One network, trained like a text model on many kinds of tasks, decides from its context whether to write words, press game buttons, or move a robot arm. The same set of weights plays Atari, captions images, chats, and stacks blocks with a real arm.
Abstract · A Generalist Agent
Inspired by progress in large-scale language modeling, we apply a similar approach towards building a single generalist agent beyond the realm of text outputs. The agent, which we refer to as Gato, works as a multi-modal, multi-task, multi-embodiment generalist policy. The same network with the same weights can play Atari, caption images, chat, stack blocks with a real robot arm and much more, deciding based on its context whether to output text, joint torques, button presses, or other tokens. In this report we describe the model and the data, and document the current capabilities of Gato.
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gomez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, et al.
arXiv:2205.06175 · cs.AI, cs.CL, cs.LG, cs.RO · submitted May 12, 2022 · updated Nov 11, 2022
abstract · pdf · html · Published at TMLR, 42 pages
the point of this paper isn't "here, we solved general intelligence". It's "look, multi modal token prediction is a sound iteration". Look at the scale of the model in comparison to, say, gpt-3: this is a PoC, they didn't bother scaling it, because we've already seen where scaling these mechanisms leads.
What I would love to know is what kind of architectures deepmind et al are playing with in-house. Token prediction is a promising avenue, but it's more of a language that an intelligent agent may operate in, opposed to the self-sufficient structure of the intelligent agent itself -- the symbolic system that implements algos like gato. If that symbolic system will be the result of a generator-function, that generator function won't be token prediction by trade. I mean, maybe somewhere in the deep depths of a multi modal model, intelligent structure may emerge, but that would be a very weird byproduct.