about
Benchmarking Neural Network Generalization for Grammar Induction (arxiv.org)
3 points by puttycat on Aug 22, 2023 | hide | past | pdf | discuss on HN

In plain words: A scoring benchmark tests how well networks learn exact formal languages, like matching counts of letters, by grading unseen strings against how much training data was used. Networks trained to compress the training data generalized better with less data than standard training.

Abstract

How well do neural networks generalize? Even for grammar induction tasks, where the target generalization is fully known, previous works have left the question open, testing very limited ranges beyond the training set and using different success criteria. We provide a measure of neural network generalization based on fully specified formal languages. Given a model and a formal grammar, the method assigns a generalization score representing how well a model generalizes to unseen samples in inverse relation to the amount of data it was trained on. The benchmark includes languages such as $a^nb^n$, $a^nb^nc^n$, $a^nb^mc^{n+m}$, and Dyck-1 and 2. We evaluate selected architectures using the benchmark and find that networks trained with a Minimum Description Length objective (MDL) generalize better and using less data than networks trained using standard loss functions. The benchmark is available at https://github.com/taucompling/bliss.

Nur Lan, Emmanuel Chemla, Roni Katzir
arXiv:2308.08253 · cs.CL · submitted Aug 16, 2023 · updated Aug 25, 2023
abstract · pdf · html · 10 pages, 4 figures, 2 tables. Conference: Learning with Small Data 2023

add comment on HN
Also discussed: Aug 2023 (2 points, 0 comments)