Application of deep learning approaches for medieval historical documents transcription

Maksym Voloshchuk; Bohdana Zarembovska; Mykola Kozlenko

arXiv:2512.18865·cs.CV·December 23, 2025

Application of deep learning approaches for medieval historical documents transcription

Maksym Voloshchuk, Bohdana Zarembovska, Mykola Kozlenko

PDF

Open Access

TL;DR

This paper develops a deep learning pipeline tailored for transcribing medieval Latin handwritten documents from the 9th to 11th centuries, addressing challenges posed by historical handwriting styles.

Contribution

It introduces a specialized dataset, analysis, and a comprehensive deep learning approach for medieval document transcription, filling a gap in historical handwriting recognition.

Findings

01

Achieved high precision and recall metrics

02

Developed a dataset for medieval Latin scripts

03

Published the implementation on GitHub

Abstract

Handwritten text recognition and optical character recognition solutions show excellent results with processing data of modern era, but efficiency drops with Latin documents of medieval times. This paper presents a deep learning method to extract text information from handwritten Latin-language documents of the 9th to 11th centuries. The approach takes into account the properties inherent in medieval documents. The paper provides a brief introduction to the field of historical document transcription, a first-sight analysis of the raw data, and the related works and studies. The paper presents the steps of dataset development for further training of the models. The explanatory data analysis of the processed data is provided as well. The paper explains the pipeline of deep learning models to extract text information from the document images, from detecting objects to word recognition…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHandwritten Text Recognition Techniques · Image Processing and 3D Reconstruction · Text and Document Classification Technologies