In plain words: Instead of judging models on one fixed train/test split, this approach re-splits the data several times across multiple datasets and estimates how likely one model truly beats the other or ties. It ranked six English part-of-speech taggers across two datasets and three scoring measures.
Abstract · Is the Best Better? Bayesian Statistical Model Comparison for Natural Language Processing
Recent work raises concerns about the use of standard splits to compare natural language processing models. We propose a Bayesian statistical model comparison technique which uses k-fold cross-validation across multiple data sets to estimate the likelihood that one model will outperform the other, or that the two will produce practically equivalent results. We use this technique to rank six English part-of-speech taggers across two data sets and three evaluation metrics.
Piotr Szymański, Kyle Gorman
arXiv:2010.03088 · cs.CL, cs.LG, stat.ME · submitted Oct 6, 2020
abstract · pdf · html · Accepted to EMNLP2020