A Tool for Facilitating OCR Postediting in Historical Documents

2020-04-23LREC 2020Code Available0· sign in to hype

Alberto Poncelas, Mohammad Aboomar, Jan Buts, James Hadley, Andy Way

Code Available — Be the first to reproduce this paper.

Code

github.com/alberto-poncelas/tesseract_postprocess
OfficialIn papernone★ 9

Abstract

Optical character recognition (OCR) for historical documents is a complex procedure subject to a unique set of material issues, including inconsistencies in typefaces and low quality scanning. Consequently, even the most sophisticated OCR engines produce errors. This paper reports on a tool built for postediting the output of Tesseract, more specifically for correcting common errors in digitized historical documents. The proposed tool suggests alternatives for word forms not found in a specified vocabulary. The assumed error is replaced by a presumably correct alternative in the post-edition based on the scores of a Language Model (LM). The tool is tested on a chapter of the book An Essay Towards Regulating the Trade and Employing the Poor of this Kingdom (Cary ,1719). As demonstrated below, the tool is successful in correcting a number of common errors. If sometimes unreliable, it is also transparent and subject to human intervention.

Tasks

Language Modeling Language Modelling Optical Character Recognition Optical Character Recognition (OCR)

A Tool for Facilitating OCR Postediting in Historical Documents

Code

Abstract

Tasks

Reproductions