Skip to main content

PDF to Markdown guide

How to preserve and repair Markdown headings from PDF

Use best-effort heading detection as a starting point and manually restore a logical heading hierarchy after extraction.

Updated September 2026 · PDFKaka EditorialOpen PDF to Markdown
Capture the PDF to Markdown workspace at 1366×768 with synthetic sample content and the controls used by this workflow visible.

Convert the PDF to Markdown

  1. Convert the stored text into the tool's best-effort headings and paragraphs.
  2. Check that headings describe the following content and do not skip levels without reason.
  3. Choose a PDF that contains selectable text, or prepare a reviewed OCR result first.

Text extraction versus OCR for scanned pages

No OCR. Uses selectable text and a heading heuristic; does not reconstruct lists, tables, images, links, or columns.

FAQ

Does PDF to Markdown OCR scanned pages?

No. Use OCR PDF first when a scanned document does not already contain usable selectable text.

What is the current input limit for PDF to Markdown?

PDF; One PDF ≤50 MiB, ≤100 pages; ≤1,000,000 extracted characters.

What does PDF to Markdown create?

selectable text extraction, best-effort reading order; no OCR; Markdown ≤150 MiB.

Open PDF to Markdown→