Text Line Segmentation of Historical Documents: a Survey
Laurence Likforman-Sulem, Abderrazak Zahour, Bruno Taconet

TL;DR
This survey reviews recent methods for automatic text line segmentation in historical documents, addressing challenges posed by low quality and artifacts, and highlights ongoing research efforts in this complex field.
Contribution
It provides a comprehensive overview of the state-of-the-art techniques for historical document text line segmentation developed over the past decade.
Findings
Various segmentation methods have been proposed for degraded historical documents.
Challenges include noise, artifacts, and interfering lines affecting segmentation accuracy.
The survey identifies gaps and future directions in the research area.
Abstract
There is a huge amount of historical documents in libraries and in various National Archives that have not been exploited electronically. Although automatic reading of complete pages remains, in most cases, a long-term objective, tasks such as word spotting, text/image alignment, authentication and extraction of specific fields are in use today. For all these tasks, a major step is document segmentation into text lines. Because of the low quality and the complexity of these documents (background noise, artifacts due to aging, interfering lines),automatic text line segmentation remains an open research field. The objective of this paper is to present a survey of existing methods, developed during the last decade, and dedicated to documents of historical interest.
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
