In plain words: A test set built from real Jupyter notebook sessions checks whether AI coding models use what the program actually printed or errored on to predict code output and write new code. Today's models do poorly, showing that using runtime information is an unexplored skill.
Abstract
In this work, we present a benchmark that consists of Jupyter notebooks development trajectories and allows measuring how large language models (LLMs) can leverage runtime information for predicting code output and code generation. We demonstrate that the current generation of LLMs performs poorly on these tasks and argue that there exists a significantly understudied domain in the development of code-based models, which involves incorporating the runtime context.
Konstantin Grotov, Sergey Titov
arXiv:2504.12365 · cs.SE, cs.AI, cs.LG · submitted Apr 16, 2025
abstract · pdf · html · Accepted to the third Deep Learning for Code (DL4C) workshop @ ICLR 2025