The Effects of OCR Error on the Extraction of Private Information [chapter]

Kazem Taghva, Russell Beckley, Jeffrey Coombs
2006 Lecture Notes in Computer Science  
OCR error has been shown not to affect the average accuracy of text retrieval or text categorization. Recent studies however have indicated that information extraction is significantly degraded by OCR error. We experimented with information extraction software on two collections, one with OCR-ed documents and another with manuallycorrected versions of the former. We discovered a significant reduction in accuracy on the OCR text versus the corrected text. The majority of errors were attributable
more » ... to zoning problems rather than OCR classification errors.
doi:10.1007/11669487_31 fatcat:xt4pnt4jsjgufkezbo7z4gmzmu