about
Vision-Language Models vs. Traditional OCR in Video – New Benchmark (arxiv.org)
6 points by ashu_trv on Feb 13, 2025 | hide | past | pdf | 1 comment on HN

In plain words: They built a free set of 1,477 hand-labeled video frames to test how well AI systems that read images can transcribe on-screen words, against older text-scanning tools. The AI systems often did better, but invented words and stumbled on blocked or fancy text.

Abstract · Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments

This paper introduces an open-source benchmark for evaluating Vision-Language Models (VLMs) on Optical Character Recognition (OCR) tasks in dynamic video environments. We present a curated dataset containing 1,477 manually annotated frames spanning diverse domains, including code editors, news broadcasts, YouTube videos, and advertisements. Three state of the art VLMs - Claude-3, Gemini-1.5, and GPT-4o are benchmarked against traditional OCR systems such as EasyOCR and RapidOCR. Evaluation metrics include Word Error Rate (WER), Character Error Rate (CER), and Accuracy. Our results highlight the strengths and limitations of VLMs in video-based OCR tasks, demonstrating their potential to outperform conventional OCR models in many scenarios. However, challenges such as hallucinations, content security policies, and sensitivity to occluded or stylized text remain. The dataset and benchmarking framework are publicly available to foster further research.

Sankalp Nagaonkar, Augustya Sharma, Ashish Choithani, Ashutosh Trivedi
arXiv:2502.06445 · cs.CV · submitted Feb 10, 2025
abstract · pdf · html · Code and dataset: https://github.com/video-db/ocr-benchmark

add comment on HN
Also discussed: Feb 2025 (142 points, 58 comments)

A new benchmark study evaluates Vision-Language Models (Claude-3, Gemini-1.5, GPT-4o) against traditional OCR tools (EasyOCR, RapidOCR) for extracting text from videos. The findings show VLMs outperforming OCR in many cases but also highlight challenges like hallucinated text and handling occluded/stylized fonts.

The dataset (1,477 manually annotated frames) and benchmarking framework are publicly available to encourage further research.

Paper: https://arxiv.org/abs/2502.06445 Dataset & Repo: https://github.com/video-db/ocr-benchmark

Would love to hear thoughts from the community on the future of VLMs in OCR.