about
From F to A on the N.Y. Regents Science Exams: An Overview of Aristo Project (arxiv.org)
2 points by godelmachine on Sep 19, 2019 | hide | past | pdf | discuss on HN

In plain words: A question-answering system built on modern language models reads eighth-grade science exam questions and picks the right answer from the choices, skipping any that need diagrams. It scored over 90% on New York's Regents exam, up from the 59.3% best score in 2016.

Abstract · From 'F' to 'A' on the N.Y. Regents Science Exams: An Overview of the Aristo Project

AI has achieved remarkable mastery over games such as Chess, Go, and Poker, and even Jeopardy, but the rich variety of standardized exams has remained a landmark challenge. Even in 2016, the best AI system achieved merely 59.3% on an 8th Grade science exam challenge. This paper reports unprecedented success on the Grade 8 New York Regents Science Exam, where for the first time a system scores more than 90% on the exam's non-diagram, multiple choice (NDMC) questions. In addition, our Aristo system, building upon the success of recent language models, exceeded 83% on the corresponding Grade 12 Science Exam NDMC questions. The results, on unseen test questions, are robust across different test years and different variations of this kind of test. They demonstrate that modern NLP methods can result in mastery on this task. While not a full solution to general question-answering (the questions are multiple choice, and the domain is restricted to 8th Grade science), it represents a significant milestone for the field.

Peter Clark, Oren Etzioni, Daniel Khashabi, Tushar Khot, Bhavana Dalvi Mishra, Kyle Richardson, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord, Niket Tandon, Sumithra Bhakthavatsalam, Dirk Groeneveld, et al.
arXiv:1909.01958 · cs.CL, cs.AI · submitted Sep 4, 2019 · updated Feb 2, 2021
abstract · pdf · html · AI Magazine 41 (4) Winter 2020. New analysis sections added

add comment on HN
Also discussed: Sep 2019 (1 point, 1 comment)