In plain words: Word boxes from a document are linked in a graph, and a network predicts which belong together, grouping boxes into lines and lines into paragraphs in one pass. It detects paragraphs as accurately as the best current systems while running far more efficiently.
Abstract · Unified Line and Paragraph Detection by Graph Convolutional Networks
We formulate the task of detecting lines and paragraphs in a document into a unified two-level clustering problem. Given a set of text detection boxes that roughly correspond to words, a text line is a cluster of boxes and a paragraph is a cluster of lines. These clusters form a two-level tree that represents a major part of the layout of a document. We use a graph convolutional network to predict the relations between text detection boxes and then build both levels of clusters from these predictions. Experimentally, we demonstrate that the unified approach can be highly efficient while still achieving state-of-the-art quality for detecting paragraphs in public benchmarks and real-world images.
Shuang Liu, Renshen Wang, Michalis Raptis, Yasuhisa Fujii
arXiv:2203.09638 · cs.CV, cs.LG · submitted Mar 17, 2022
abstract · pdf · html · Accepted to DAS 2022 as an oral paper
It was surfaced in iOS a decade ago as "tap to zoom" feature for PDFs. It's funny — as with a lot of things there was a lot of sophisticated engineering under the hood and then marketing simply wants it to detect a tap in a paragraph and zoom to its bounds.
I can't think of the last time I read a PDF on my phone or I would test it to see if it still works as I remember.