About this tool
Turn scanned pages into editable English text with signature checks, explicit PDF rendering and page-layout policies, engine confidence, resource limits, and source-review evidence.
Scanned PDF & Image OCR converts the pixels of a scan into editable English text using Tesseract.js 6.0.1, with the recognition engine running in your browser. It accepts a PDF or a PNG, JPEG, WebP, GIF, or AVIF image up to 25 MB, checks the file signature and real dimensions before starting, and for PDFs renders up to 10 selected pages through PDF.js at a resolution you choose. The output is a text area you can correct in place and download as UTF-8 TXT, plus a per-page report of engine confidence, word and character counts, and render dimensions. Its purpose is turning a static scan into a draft transcript you can search, quote, or paste elsewhere, not producing a final, authoritative copy.
- Checks PDF, PNG, JPEG, WebP, GIF, or AVIF signatures and real image dimensions before loading the Tesseract.js 6.0.1 English engine.
- Selects up to 10 PDF pages, renders at 1.5x, 2x, or 2.5x, and enforces pixel, prepared-data, and recognized-output ceilings with immediate PDF.js and OCR worker cancellation.
- Reports per-page confidence, word and character counts, empty results, render dimensions, recognition settings, and whether exported text was edited after OCR.
How to use PDF & Image OCR
Choose or drop a PDF or image, and the source panel shows what was detected. For PDFs, type a page selection such as 1-3, 5 in the Pages field (blank means all, up to 10), pick a PDF render resolution of 1.5x, 2x, or 2.5x, and optionally tick Include PDF page separators to get a header line per page. Set Page layout to Automatic layout for ordinary documents, Single text block for a clean paragraph crop, or Sparse text for scattered labels. Press Run OCR; the first run downloads the English engine assets, and Cancel OCR stops rendering or recognition immediately. Edit the recognized text directly, then click Download TXT. The report notes whether the text was edited after recognition.
When this tool is useful
- A researcher needs quotable text from a scanned journal article that arrived as an image-only PDF.
- An accounts clerk wants to search a stack of scanned invoices for a vendor name instead of reading each one.
- A student photographs a printed handout with a phone and needs the text in their notes app.
- A translator receives a JPEG screenshot of English copy and needs an editable draft to work from.
- A records team extracts typed text from old memos before deciding whether a full transcription is worth commissioning.
Practical tips
- Start at 2x for typical 300 dpi scans. Drop to 1.5x if you hit the 80-megapixel batch limit, and use 2.5x only for small print.
- Crop, straighten, and boost contrast in an image editor first; skewed or low-contrast scans lower confidence more than any setting here can recover.
- A mean confidence around 90% still means roughly one word in ten deserves a second look, so read names, dates, and amounts against the source.
- Handwriting, tables, multi-column layouts, and non-English text are outside what this English-only engine handles reliably.
- Any change to pages, resolution, or layout marks the result stale; re-run OCR before downloading or the TXT will not match the settings shown.
Examples you can test
Load an example, compare the result with the expected output, then replace it with your own input.
Extract two pages from a scanned contract
Example input
contract.pdf, 14 image-only pages; Pages set to 3, 7; resolution 2x; separators on
Expected output
Text for pages 3 and 7 with a separator header before each, plus per-page confidence such as 91.4% and 88.7%
Page numbers in the report refer to the PDF's physical page order, so page 7 is the seventh page even if the printed footer says otherwise.
Read a product label photo
Example input
label.jpg, 3000x2000 px, with a few short lines of text on a plain background; Page layout set to Sparse text
Expected output
The label lines as separate paragraphs and a single-page confidence figure
Sparse text stops the engine from trying to merge distant fragments into one block, which usually improves short-label results.
Validation checklist
- Compare the page count in the report with the pages you selected.
- Read every proper noun, number, and date against the original scan.
- Check that paragraph breaks and list order survived, especially in two-column pages.
- Confirm the downloaded TXT reflects your manual corrections, not the raw engine output.
- Keep the source file beside the text until someone has reviewed it.