Thank you! Thank you so much for helping to correct thousands, if not millions, of pages of scanned text! What? You aren't doing anything of the sort? Oh, yes you are ... read on! Optical character recognition (OCR) is the method used to digitise printed material — old newspapers and book, for example, can be scanned using OCR technology, and then we can access them online. The National Library of Australia uses this technology for its massive Trove database, for example. The problem with OCR is that it's not always that accurate, especially from older printed material with yellowing paper and faded or smudged ink. It is a lot better than it used to be, but is still far from 100% accurate — 80% accuracy is more typical. So the resulting scanned texts have a lot of errors! Here's an example: We can read the top sentence of scanned type (This aged portion of society were distinguished from ...) , because us humans are fucking brilliant, but a computer h...
A blog for people who love puzzles. With a little indexing and editing on the side.