Add searchable text layer to scanned PDF documents using OCR
OCR PDF
Add searchable text layer to scanned PDF documents using OCR
CHOOSE FILES
(or drag them here)
The tool runs Tesseract OCR server-side. It adds a searchable text layer over the existing scanned image without removing or replacing the original scan — the visual appearance of each page is unchanged. The text becomes selectable and Ctrl+F searchable. If you later want to extract the text to a plain .txt file, use the PDF to Text tool on the OCR’d output.
When OCR works well and when it doesn’t
Clean printed text on a flat, well-lit scan at 300 DPI or above: excellent accuracy. Handwriting: patchy at best. Mixed-orientation pages (some portrait, some landscape): the engine handles rotation before recognition, so this is usually fine. Very low contrast, coffee-stained, or heavily wrinkled documents: expect errors. The Enhanced mode applies pre-processing steps that help with marginal scans but won’t rescue a truly unreadable page.
Language support
Over 100 languages from a dropdown, including Arabic, Chinese (Simplified and Traditional), Japanese, Korean, Hindi, and all major European scripts. Select the language that matches the document’s primary script for best accuracy. Mixed-script documents — like a bilingual contract — can be processed but require the dominant language as the selection; the secondary script will get partial recognition.
OCR modes
- Standard
- Fast processing; reliable for clear, printed text. Most documents don’t need anything more.
- Enhanced
- Slower; applies additional pre-processing that improves recognition on low-resolution scans, faded ink, or pages with skew.
