about
A Flawed Dataset for Symbolic Equation Verification (arxiv.org)
2 points by YeGoblynQueenne on Mar 6, 2022 | hide | past | pdf | discuss on HN

In plain words: A proposed way to build practice data for checking symbolic equations is flawed: its true equations cover only a narrow slice, and its fake ones are made differently so they are easy to spot. A simple guessing rule already solves the task almost perfectly.

Abstract

Arabshahi, Singh, and Anandkumar (2018) propose a method for creating a dataset of symbolic mathematical equations for the tasks of symbolic equation verification and equation completion. Unfortunately, a dataset constructed using the method they propose will suffer from two serious flaws. First, the class of true equations that the procedure can generate will be very limited. Second, because true and false equations are generated in completely different ways, there are likely to be artifactual features that allow easy discrimination. Moreover, over the class of equations they consider, there is an extremely simple probabilistic procedure that solves the problem of equation verification with extremely high reliability. The usefulness of this problem in general as a testbed for AI systems is therefore doubtful.

Ernest Davis
arXiv:2105.11479 · cs.AI, cs.SC · submitted May 24, 2021 · updated May 28, 2021
abstract · pdf · html

add comment on HN