In plain words: General-purpose chat AI models were tested on pulling structured details from Spanish electricity bills without any special training, varying the prompt wording and settings. Adding worked examples to the prompt mattered far more than tuning settings, lifting its correctness score by over 19 points to 97.6%.
Abstract · Information Extraction from Electricity Invoices with General-Purpose Large Language Models
Information extraction from semi-structured business documents remains a critical challenge for enterprise management. This study evaluates the capability of general-purpose Large Language Models to extract structured information from Spanish electricity invoices without task-specific fine-tuning. Using a subset of the IDSEM dataset, we benchmark two architecturally distinct models, Gemini 1.5 Pro and Mistral-small, across 19 parameter configurations and 6 prompting strategies. Our experimental framework treats prompt engineering as the primary experimental variable, comparing zero-shot baselines against increasingly sophisticated few-shot approaches and iterative extraction strategies. Results demonstrate that prompt quality dominates over hyperparameter tuning: the F1-score variation across all parameter configurations is marginal, while the gap between zero-shot and the best few-shot strategy exceeds 19 percentage points. The best configuration (few-shot with cross-validation) achieves an F1-score of 97.61% for Gemini and 96.11% for Mistral-small, with document template structure emerging as the primary determinant of extraction difficulty. These findings establish that prompt design is the critical lever for maximizing extraction fidelity in LLM-based document processing, thereby providing an empirical framework for integrating general-purpose LLMs into business document automation.
Javier Gómez, Javier Sánchez
arXiv:2604.25927 · cs.CL · submitted Apr 1, 2026
abstract · pdf · html · 13 pages, 2 figures