OCR PDF guide
Why OCR misses words, names, or numbers in a PDF
OCR errors usually begin in the page image: the software can only interpret the shapes, contrast, orientation, and language information that reach it.

On this page
Image quality determines what OCR can see
Low-resolution characters contain fewer distinguishing details. Blur joins strokes together, shadows hide edges, and weak contrast makes letters blend into the page. A tilted or perspective-distorted photograph also changes the shapes the recognizer expects.
Whenever possible, recapture the page in even light, keep the camera parallel to the paper, fill the frame without cutting off margins, and make sure the smallest important text is readable before creating the PDF.
Names and numbers need extra review
OCR uses language patterns to choose between similar shapes. Ordinary words provide context; names, reference codes, dates, and account numbers often do not. A zero can resemble the letter O, a one can resemble lowercase l, and punctuation can disappear beside noisy backgrounds.
Compare these values character by character with the scan. Do not rely on search alone to validate a legal name, financial amount, medicine dose, or safety-critical number.
Layout can scramble otherwise accurate words
Columns, tables, sidebars, stamps, and rotated labels create several possible reading orders. OCR may recognize the words correctly while placing them in the wrong sequence. Tables can lose row and column relationships, and page headers may appear inside the main text.
Review sample passages from every layout type in the document rather than checking only the first page.
Language and handwriting boundaries
PDFKaka's current OCR is intended for printed English. Handwriting and other languages are not supported as reliable workflows. Decorative lettering, curved text, and heavily stylized fonts may also behave more like images than ordinary printed text.
Improve the source before repeating OCR
- Rotate pages upright and replace blurred captures where possible.
- Improve lighting and contrast without erasing thin character strokes.
- Split a mixed-quality document and test a representative page first.
- Run OCR again on the improved source.
- Search for common words, then manually verify names, numbers, and structured content.
FAQ
Why are names and numbers often misread?
They provide less language context than ordinary sentences, and characters such as O/0 or l/1 have similar shapes. Check them directly against the scan.
Does PDFKaka OCR handwriting?
The current workflow is intended for printed English. Handwriting is not a supported reliable use case.
Why does skewed text reduce accuracy?
Skew and perspective distortion change character shapes and line alignment, making segmentation and recognition harder.
How many scanned pages can one OCR job contain?
The verified OCR limit is 25 scanned pages, within the general PDF input limits.