about
Transformers Utilization in Chart Understanding: A Review of Advances and Future (arxiv.org)
39 points by sandwichsphinx on Oct 22, 2024 | hide | past | pdf | 2 comments on HN

In plain words: This review sorts 32 studies from 2020–2024 on AI systems that read charts by linking the image, its data table, and the question in one pass. They beat the old rule-based tricks, but still stumble on blurry images, text reading, and chart reasoning.

Abstract · Transformers Utilization in Chart Understanding: A Review of Recent Advances & Future Trends

In recent years, interest in vision-language tasks has grown, especially those involving chart interactions. These tasks are inherently multimodal, requiring models to process chart images, accompanying text, underlying data tables, and often user queries. Traditionally, Chart Understanding (CU) relied on heuristics and rule-based systems. However, recent advancements that have integrated transformer architectures significantly improved performance. This paper reviews prominent research in CU, focusing on State-of-The-Art (SoTA) frameworks that employ transformers within End-to-End (E2E) solutions. Relevant benchmarking datasets and evaluation techniques are analyzed. Additionally, this article identifies key challenges and outlines promising future directions for advancing CU solutions. Following the PRISMA guidelines, a comprehensive literature search is conducted across Google Scholar, focusing on publications from Jan'20 to Jun'24. After rigorous screening and quality assessment, 32 studies are selected for in-depth analysis. The CU tasks are categorized into a three-layered paradigm based on the cognitive task required. Recent advancements in the frameworks addressing various CU tasks are also reviewed. Frameworks are categorized into single-task or multi-task based on the number of tasks solvable by the E2E solution. Within multi-task frameworks, pre-trained and prompt-engineering-based techniques are explored. This review overviews leading architectures, datasets, and pre-training tasks. Despite significant progress, challenges remain in OCR dependency, handling low-resolution images, and enhancing visual reasoning. Future directions include addressing these challenges, developing robust benchmarks, and optimizing model efficiency. Additionally, integrating explainable AI techniques and exploring the balance between real and synthetic data are crucial for advancing CU research.

Mirna Al-Shetairy, Hanan Hindy, Dina Khattab, Mostafa M. Aref
arXiv:2410.13883 · cs.CV, cs.AI, cs.HC, cs.LG · submitted Oct 5, 2024
abstract · pdf · html

add comment on HN

This is a critical space for progressing AI in science right now. Once we have algorithms that can process charts and interpret data, our ability to integrate scientific information from multiple studies will increase exponentially. We may even find new interpretations from charted data that human eyes are unable to interpret. Until then, it's as if we're stuck with blind AI agents who have to take text for granted without validating it against graphical data.
Have they automated PRISMA lit reviews yet?