In plain words: This survey maps how machines read and understand business documents, tracing the field from hand-written rules and statistical methods to deep learning models trained on many examples. It lays out the main tasks, models, and test sets, and points to open questions ahead.
Abstract · Document AI: Benchmarks, Models and Applications
Document AI, or Document Intelligence, is a relatively new research topic that refers to the techniques for automatically reading, understanding, and analyzing business documents. It is an important research direction for natural language processing and computer vision. In recent years, the popularity of deep learning technology has greatly advanced the development of Document AI, such as document layout analysis, visual information extraction, document visual question answering, document image classification, etc. This paper briefly reviews some of the representative models, tasks, and benchmark datasets. Furthermore, we also introduce early-stage heuristic rule-based document analysis, statistical machine learning algorithms, and deep learning approaches especially pre-training methods. Finally, we look into future directions for Document AI research.
Lei Cui, Yiheng Xu, Tengchao Lv, Furu Wei
arXiv:2111.08609 · cs.CL · submitted Nov 16, 2021
abstract · pdf · html