about
SmolDocling: An ultra-compact VLM for end-to-end multi-modal document conversion (arxiv.org)
66 points by prats226 on Mar 21, 2025 | hide | past | pdf | 12 comments on HN

In plain words: A tiny model reads a whole document page and writes one markup format that records every element's text, layout, and position, replacing pipelines of separate specialized tools. It matches models up to 27 times larger while using far less computing power.

Abstract · SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

We introduce SmolDocling, an ultra-compact vision-language model targeting end-to-end document conversion. Our model comprehensively processes entire pages by generating DocTags, a new universal markup format that captures all page elements in their full context with location. Unlike existing approaches that rely on large foundational models, or ensemble solutions that rely on handcrafted pipelines of multiple specialized models, SmolDocling offers an end-to-end conversion for accurately capturing content, structure and spatial location of document elements in a 256M parameters vision-language model. SmolDocling exhibits robust performance in correctly reproducing document features such as code listings, tables, equations, charts, lists, and more across a diverse range of document types including business documents, academic papers, technical reports, patents, and forms -- significantly extending beyond the commonly observed focus on scientific papers. Additionally, we contribute novel publicly sourced datasets for charts, tables, equations, and code recognition. Experimental results demonstrate that SmolDocling competes with other Vision Language Models that are up to 27 times larger in size, while reducing computational requirements substantially. The model is currently available, datasets will be publicly available soon.

Ahmed Nassar, Andres Marafioti, Matteo Omenetti, Maksym Lysak, Nikolaos Livathinos, Christoph Auer, Lucas Morin, Rafael Teixeira de Lima, Yusik Kim, A. Said Gurbuz, Michele Dolfi, Miquel Farré, et al.
arXiv:2503.11576 · cs.CV · submitted Mar 14, 2025
abstract · pdf · html · 24 pages, 10 figures

add comment on HN

After many posts on my feed, I decided to give it a spin.

The good: - Open source.

- Can run locally (Apple Silicon) at a fair speed.

- Image detection is good.

The bad:

- Not detecting tables.

- Text in a perfectly clean PDF (resume) is not detected.

I know its in preview, small and open source which is great, but its far from being usable.

-

Well it's certainly small. Absolutely bombs my KTANE test though - poor character recognition, poor handling of even mildly complex tables, and prone to getting stuck in repetition loops. (Task was convert to docling, in the official HF space.)

That said, I'm definitely glad to see work in this area, particularly with open weights.

https://news.ycombinator.com/item?id=43431609 mentioned that it was fine tuned from https://huggingface.co/HuggingFaceTB/SmolVLM-256M-Instruct

It would be interesting to see if fine tuning on your KTANE test improves your results?

I may be missing something, but if you train a model to pass a specific test… isn’t it obvious that it would do better on that test?

I thought we called models with test data in their training set “poisoned”

That is how training is done. You don't train on all of your material or the tests are worthless.
If you train on any test material, the tests are worthless.
Regarding the repetition loops, I found that adding the end of turn token to the stop param was enough. Documentation mentions detecting this.

But your point about quality stands. Separately, this model emits the docling XML format, not the JSON format, so as far as I know today that means you are using the Python flavored docling only, the JS variant does not support this yet (afaik).

It also can be used here:https://www.smoldocling.net, works well!
What’s the best library for fine-tuning VLMs at the moment and do they support this architecture or that for the IBM Granite vision models? Document understanding tasks seem in special need of fine-tuning.
It looks like the model itself is here: https://huggingface.co/ds4sd/SmolDocling-256M-preview

It was fine tuned from this: https://huggingface.co/HuggingFaceTB/SmolVLM-256M-Instruct

There's an example of fine tuning the base that would likely be applicable to this one as well.

Does seem comparable to Tesseract? I feel like the accuracy results are still not significantly improved as a whole.
OCR is not the task being solved here, though. This is supposed to help you when dealing with complex layouts where text is not just read left-to-right, top-to-bottom.

But I agree that accurate OCR is kind of a prerequisite for adaptation.