Recent advancements in Large Language Models (LLMs) have opened new opportunities for automating complex tasks, including Handwritten Text Recognition (HTR) and named entity extraction from historical manuscripts. This paper presents an approach that leverages LLMs to perform HTR directly on handwritten documents, reducing the reliance on traditional handwriting recognition tools. The system enhances transcription accuracy by incorporating repetitive text patterns and a database of names and surnames. In addition to HTR, the study explores using LLMs for extracting named entities (NEs) from structured and highly repetitive texts, such as historical population registries. The approach evaluates the performance of GPT-3.5 Turbo and GPT-4 in NE extraction tasks, comparing the effects of different levels of instruction detail. For HTR, the impact of integrating an NE database is assessed against a baseline model without additional data support. The results demonstrate that providing LLMs with structured guidance and external NE databases significantly improves both transcription quality and entity recognition accuracy. While GPT-4 achieves the highest precision, GPT-3.5 Turbo presents a cost-effective alternative with a good balance between performance and efficiency. These findings highlight the potential of LLMs as versatile tools for both manuscript transcription and structured data extraction from historical records.
Using LLMs for Semi-automatic Handwritten Text Recognition and Elaboration
Angelica Lo Duca
2026
Abstract
Recent advancements in Large Language Models (LLMs) have opened new opportunities for automating complex tasks, including Handwritten Text Recognition (HTR) and named entity extraction from historical manuscripts. This paper presents an approach that leverages LLMs to perform HTR directly on handwritten documents, reducing the reliance on traditional handwriting recognition tools. The system enhances transcription accuracy by incorporating repetitive text patterns and a database of names and surnames. In addition to HTR, the study explores using LLMs for extracting named entities (NEs) from structured and highly repetitive texts, such as historical population registries. The approach evaluates the performance of GPT-3.5 Turbo and GPT-4 in NE extraction tasks, comparing the effects of different levels of instruction detail. For HTR, the impact of integrating an NE database is assessed against a baseline model without additional data support. The results demonstrate that providing LLMs with structured guidance and external NE databases significantly improves both transcription quality and entity recognition accuracy. While GPT-4 achieves the highest precision, GPT-3.5 Turbo presents a cost-effective alternative with a good balance between performance and efficiency. These findings highlight the potential of LLMs as versatile tools for both manuscript transcription and structured data extraction from historical records.| File | Dimensione | Formato | |
|---|---|---|---|
|
LoDuca_WEBIST2024.pdf
accesso aperto
Tipologia:
Documento in Post-print
Licenza:
Creative commons
Dimensione
4.19 MB
Formato
Adobe PDF
|
4.19 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


