Skip to main content

OCR PDF guide

Why OCR misses words, names, or numbers in a PDF

OCR errors usually begin in the page image: the software can only interpret the shapes, contrast, orientation, and language information that reach it.

Updated September 2026 · PDFKaka EditorialOpen OCR PDF
Show the relevant OCR PDF result, warning, or diagnostic state with synthetic sample content; include the clue the article tells the reader to inspect.
On this page
  1. Image quality determines what OCR can see
  2. Names and numbers need extra review
  3. Layout can scramble otherwise accurate words
  4. Language and handwriting boundaries
  5. Improve the source before repeating OCR

Image quality determines what OCR can see

Low-resolution characters contain fewer distinguishing details. Blur joins strokes together, shadows hide edges, and weak contrast makes letters blend into the page. A tilted or perspective-distorted photograph also changes the shapes the recognizer expects.

Whenever possible, recapture the page in even light, keep the camera parallel to the paper, fill the frame without cutting off margins, and make sure the smallest important text is readable before creating the PDF.

Names and numbers need extra review

OCR uses language patterns to choose between similar shapes. Ordinary words provide context; names, reference codes, dates, and account numbers often do not. A zero can resemble the letter O, a one can resemble lowercase l, and punctuation can disappear beside noisy backgrounds.

Compare these values character by character with the scan. Do not rely on search alone to validate a legal name, financial amount, medicine dose, or safety-critical number.

Layout can scramble otherwise accurate words

Columns, tables, sidebars, stamps, and rotated labels create several possible reading orders. OCR may recognize the words correctly while placing them in the wrong sequence. Tables can lose row and column relationships, and page headers may appear inside the main text.

Review sample passages from every layout type in the document rather than checking only the first page.

Language and handwriting boundaries

PDFKaka's current OCR is intended for printed English. Handwriting and other languages are not supported as reliable workflows. Decorative lettering, curved text, and heavily stylized fonts may also behave more like images than ordinary printed text.

Improve the source before repeating OCR

  1. Rotate pages upright and replace blurred captures where possible.
  2. Improve lighting and contrast without erasing thin character strokes.
  3. Split a mixed-quality document and test a representative page first.
  4. Run OCR again on the improved source.
  5. Search for common words, then manually verify names, numbers, and structured content.

FAQ

Why are names and numbers often misread?

They provide less language context than ordinary sentences, and characters such as O/0 or l/1 have similar shapes. Check them directly against the scan.

Does PDFKaka OCR handwriting?

The current workflow is intended for printed English. Handwriting is not a supported reliable use case.

Why does skewed text reduce accuracy?

Skew and perspective distortion change character shapes and line alignment, making segmentation and recognition harder.

How many scanned pages can one OCR job contain?

The verified OCR limit is 25 scanned pages, within the general PDF input limits.

Open OCR PDF→