about
Self-Supervised Behavior Cloned Transformers Are Path Crawlers for Text Games (arxiv.org)
19 points by PaulHoule on Dec 16, 2023 | hide | past | pdf | 2 comments on HN

In plain words: A game-playing model trains itself by exploring move sequences that reach a reward, then quickly testing small copies on new games to keep only the paths that generalize. It reaches about 90% of the performance of models trained on hand-labeled play across three text games.

Abstract · Self-Supervised Behavior Cloned Transformers are Path Crawlers for Text Games

In this work, we introduce a self-supervised behavior cloning transformer for text games, which are challenging benchmarks for multi-step reasoning in virtual environments. Traditionally, Behavior Cloning Transformers excel in such tasks but rely on supervised training data. Our approach auto-generates training data by exploring trajectories (defined by common macro-action sequences) that lead to reward within the games, while determining the generality and utility of these trajectories by rapidly training small models then evaluating their performance on unseen development games. Through empirical analysis, we show our method consistently uncovers generalizable training data, achieving about 90\% performance of supervised systems across three benchmark text games.

Ruoyao Wang, Peter Jansen
arXiv:2312.04657 · cs.CL, cs.AI · submitted Dec 7, 2023
abstract · pdf · html · Accepted to EMNLP 2023 (Findings)

add comment on HN

On one hand this describes the way forward, doesn't it? In order for transformers to truly beat humans at tasks, they need to be free to try these tasks in the real world.

If you do this in a game world, they can explore the game world. But what happens if you let them loose in the real world?

Is there a difference? Ie imagine a robot with only one video feed. No other sensors. Would a realistic game world be any different than the real world?

Our current tech falls short for simulations I’m sure, but I suspect the differences between real and simulation will revolve around the inputs the machine has available. Thoughts?