about
Meta: Gaia - A Benchmark for General AI Assistants (arxiv.org)
36 points by georgehill on Nov 23, 2023 | hide | past | pdf | 8 comments on HN

In plain words: A test set of 466 real-world questions that are easy for people but demand reasoning, reading images, web browsing, and tool use to answer. Humans scored 92% on them, while GPT-4 with plugins managed just 15%.

Abstract · GAIA: a benchmark for General AI Assistants

We introduce GAIA, a benchmark for General AI Assistants that, if solved, would represent a milestone in AI research. GAIA proposes real-world questions that require a set of fundamental abilities such as reasoning, multi-modality handling, web browsing, and generally tool-use proficiency. GAIA questions are conceptually simple for humans yet challenging for most advanced AIs: we show that human respondents obtain 92\% vs. 15\% for GPT-4 equipped with plugins. This notable performance disparity contrasts with the recent trend of LLMs outperforming humans on tasks requiring professional skills in e.g. law or chemistry. GAIA's philosophy departs from the current trend in AI benchmarks suggesting to target tasks that are ever more difficult for humans. We posit that the advent of Artificial General Intelligence (AGI) hinges on a system's capability to exhibit similar robustness as the average human does on such questions. Using GAIA's methodology, we devise 466 questions and their answer. We release our questions while retaining answers to 300 of them to power a leader-board available at https://huggingface.co/gaia-benchmark.

Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, Thomas Scialom
arXiv:2311.12983 · cs.CL, cs.AI · submitted Nov 21, 2023
abstract · pdf · html

add comment on HN

Here are some example questions from the paper[0]

Level 1 Question: What was the actual enrollment count of the clinical trial on H. pylori in acne vulgaris patients from Jan-May 2018 as listed on the NIH website? Ground truth: 90

Level 2 <photo of ice cream container showing nutrition facts> Question: If this whole pint is made up of ice cream, how many percent above or below the US federal standards for butterfat content is it when using the standards as reported by Wikipedia in 2020? Answer as + or - a number rounded to one decimal place. Ground truth: +4.6

Level 3 Question: In NASA’s Astronomy Picture of the Day on 2006 January 21, two astronauts are visible, with one appearing much smaller than the other. As of August 2023, out of the astronauts in the NASA Astronaut Group that the smaller astronaut was a member of, which one spent the least time in space, and how many minutes did he spend in space, rounded to the nearest minute? Exclude any astronauts who did not spend any time in space. Give the last name of the astronaut, separated from the number of minutes by a semicolon. Use commas as thousands separators in the number of minutes. Ground truth: White; 5876

[0]: https://arxiv.org/pdf/2311.12983.pdf

The paper says, "Solving GAIA requires full automation since no approximation is allowed in the answer." While they tested GPT-4 Turbo with addons, they note they couldn't test multi-modal and addons at the same time, which lowered the scores. I surmise a properly-crafted GPT might score much higher than 15%.
Hi some authors of the work here, thanks a lot for sharing the paper, it's been quite some time in the work and we're super happy to share it with the world.

A short note on some of the reasons we decided to go with openly-sharing the questions instead of holding them back (which was another option we contemplated): - with closed-models we need to send the questions through an external AI anyway so a full privacy of the test set is not possible in general unless the leaderboard is restricted to open models (would be quite restrictive) - also, the benchmark contain a limited number of questions which are non-obvious and take a significant time for human reviewers to solve. We thus don't expect the dataset to become training material for models and to lead to having model over-fitting on the benchmark pattern in the traditional sense that happened with larger benchmark datasets including a training split. This benchmark is generally closer in philosophy to small, hand crafted benchmark datasets, like HumanEval for instance has been for code models.

> We release our questions while retaining answers to 300 of them to power a leader-board available at this https URL.

Nothing on the leaderboard: https://huggingface.co/gaia-benchmark

Good, but honestly they should have retained the questions too!

And yeah, the leaderboard appears to be down with a nice Python stacktrace.

We need to double check this questions manually first if the answers are actually correct
First author here, half of the crafting process actually consists in double checking the questions
interesting take. another milestone will be achieved when GAIA defeated. who will be the first? :)