Home  ›  Knowledge Centre  ›  What Is OCR Scanning?
Scanned document with OCR text recognition
KNOWLEDGE CENTRE · OCR & DIGITAL DOCUMENTS

What Is OCR Scanning?

Understanding optical character recognition, searchable PDFs, OCR accuracy and why the image remains important.

Understanding optical character recognition, searchable PDFs, OCR accuracy and why the image remains important.

What is OCR?

OCR stands for Optical Character Recognition. It analyses a digital image of text and attempts to convert visible characters into machine-readable text that can be searched, copied or indexed.

What is an OCR PDF?

An OCR PDF can contain the original page image together with a hidden text layer. Users can search the document while still seeing the page as captured.

Stage 1 — Capture a good image

OCR quality begins with image quality. Resolution, focus, contrast, page alignment and colour information can all affect recognition.

Stage 2 — Recognise the text

OCR software analyses the page and identifies likely characters, words and layout. Printed modern text is generally easier to recognise than poor-quality type, handwriting, historical fonts or degraded carbon copies.

Stage 3 — Quality assurance

OCR should be tested rather than assumed to be perfect. A project can sample pages and search for common errors, missing text and problems caused by unusual layouts.

Stage 4 — Preserve the image

The OCR text layer should not be treated as a replacement for the original image. The captured page remains the visual reference and can be used to verify questionable OCR results.

OCR for archives

OCR can make large document collections much more useful by enabling full-text searching. For historical collections, accuracy depends heavily on the source material and may require realistic expectations or manual correction.

The practical approach

The best OCR workflow starts with a good scan, uses appropriate recognition settings and includes quality checks appropriate to the purpose of the collection.

Frequently Asked Questions

Can handwriting be OCR scanned?

Handwriting recognition is possible in some circumstances but is considerably more dependent on handwriting style, image quality and software than conventional printed OCR.

Can old documents be OCR'd?

Yes, but historical typefaces, faded print, bleed-through and damage can reduce accuracy.

Should OCR replace TIFF masters?

No. OCR-derived text is an access and discovery layer; high-quality page images should remain available as the visual source.

Further Reading

Oxford Duplication Centre — Practical Expertise

Our approach to heritage digitisation is based on matching the capture method to the physical material. Assessment, safe handling, appropriate equipment, consistent image capture and quality assurance are all part of the process.

Need DAT tapes digitised?

Oxford Duplication Centre provides professional DAT digitisation for individual tapes and archive collections, with preservation and access outputs tailored to your requirements.

Document Scanning & OCR Request a Quote

Continue exploring the Knowledge Centre