Skip to main content

Redact PDF guide

Why Scanned PDFs Need OCR Before Text-Based Redaction

A scanned PDF stores words as pixels, so PDFKaka's literal selectable-text matcher cannot directly find arbitrary text inside those page images.

Updated September 2026 · PDFKaka EditorialCheck the scan with OCR PDF
Create an original diagram for “Why Scanned PDFs Need OCR Before Text-Based Redaction” using the verified input, processing, and output boundaries for Redact PDF.

Why the normal redaction search does not see a scan

Redact PDF searches literal substrings within individual selectable-text runs. A photographed or scanned page may contain no text runs at all. The tool also cannot draw arbitrary redaction rectangles over image regions.

What OCR can and cannot change

OCR PDF can add a best-effort printed-English text layer to suitable scans. That can make some words searchable, but recognition may miss characters, split one phrase across text runs, or place text boxes imperfectly. Handwriting and other languages are not supported reliable workflows.

After OCR, Redact PDF can only act on literal matches that its text-run search actually finds. It rasterizes output pages, but a missed match remains a missed match.

When to stop this workflow

Do not use this path for high-risk scanned redaction when the required content cannot be matched and verified precisely. Choose a redaction workflow that supports deliberate area selection and independent review. PDFKaka should not claim completion when its current controls cannot safely express the task.

FAQ

Can PDFKaka draw a redaction box over part of a scanned page?

No. The current Redact PDF tool does not provide arbitrary box redaction.

Will OCR make every scanned word safe to redact by search?

No. OCR is best effort, and Redact PDF only covers literal matches it finds within individual text runs. Every sensitive item must be checked.

Does Redact PDF recognize handwriting in scans?

No reliable handwriting workflow is supported. Current OCR is intended for printed English.

Check the scan with OCR PDF→