about
Radically Lower Data-Labeling Costs for Visual Rich Document Extraction Models (arxiv.org)
2 points by PaulHoule on Nov 1, 2022 | hide | past | pdf | discuss on HN

In plain words: Instead of typing every field on thousands of invoices, labelers just answer yes or no to answers a model already guessed, with the trickiest guesses sent first. Across three document types, this cut labeling costs 10 times with barely any drop in accuracy.

Abstract · Radically Lower Data-Labeling Costs for Visually Rich Document Extraction Models

A key bottleneck in building automatic extraction models for visually rich documents like invoices is the cost of acquiring the several thousand high-quality labeled documents that are needed to train a model with acceptable accuracy. We propose Selective Labeling to simplify the labeling task to provide "yes/no" labels for candidate extractions predicted by a model trained on partially labeled documents. We combine this with a custom active learning strategy to find the predictions that the model is most uncertain about. We show through experiments on document types drawn from 3 different domains that selective labeling can reduce the cost of acquiring labeled data by $10\times$ with a negligible loss in accuracy.

Yichao Zhou, James B. Wendt, Navneet Potti, Jing Xie, Sandeep Tata
arXiv:2210.16391 · cs.CL · submitted Oct 28, 2022
abstract · pdf · html · 9 pages, 8 figures, 3 tables

add comment on HN