Confidence-Based End-to-End Extraction of Large Historical Economic Tables with LLMs

Referierte Aufsätze Web of Science

Luzian Uihlein, Norbert Fischer, Alexander Hartelt, Tobias Gebel, Frank Puppe

In: Applied Sciences 16 (2026), 14, 7327, 18 S.

Abstract

Optical document recognition has made great progress in recent years. However, large historical tables still pose substantial challenges because of poor scan quality, irregular layouts, and the sheer number of table cells. Table transcription tools perform well on small tables but struggle with large ones. Most publications report transcription quality on benchmark data sets. From a practical perspective, this is not sufficient: when an application requires a manual post-correction phase, the required effort depends primarily on the fraction of the table that must be reviewed, and less on precision and recall values alone. We therefore introduce an additional metric: the fraction of the table that does not require manual validation because it has been transcribed with high confidence. We present two confidence-aware approaches for transcribing large historical tables: a conventional neural network pipeline and an approach based on large language models (LLMs). In our experiments, multimodal LLMs available before late 2025 did not reliably preserve the structure and content of full-page historical tables of this size and complexity; recent models, however, now outperform conventional neural networks even without domain-specific training. With several enhancements, an ensemble of three LLMs achieved a high-confidence share of 99.5% for data cells, making manual validation unnecessary for these cells. We evaluated the high-confidence data cells and found a data cell error rate of only 0.015%. Our data set contains 103 large historical German economic tables with complex structures, comprising approximately 86,000 data cells (numbers) and approximately 30,000 row header cells.

Tobias Gebel

Deputy Head of Department IT and Research Infrastructure Department



Keywords: table transcription, confidence-based reasoning, LLM, OCR, layout recognition
DOI:
https://doi.org/10.3390/app16147327

keyboard_arrow_up