Convert documents

Make a scanned PDF searchable

Make scanned pages searchable in 100+ languages while keeping each page image.

Choose a file

Choose one scanned PDF, up to 80 pages or 120 MB.

or drag and drop hereRuns on this deviceNo file upload

Make a scanned PDF searchable

OCR adds an invisible text layer to scanned pages so you can search, select and copy text while the page image remains visible.

The page image stays visible

A searchable PDF keeps the image of each page and adds an invisible text layer behind it. The page remains visually unchanged while search, text selection, copying and indexing become available. OCR workflows that re-render a page can change its sharpness or colour while adding text.

The source page images are copied into the output without being re-rendered. Recognised text is added line by line at the position where it was found, leaving the visible page image in place.

How text recognition works

Each page that needs recognition is rendered at up to 3,200 pixels on its long edge, high enough that small print, such as eight-point footers, keeps its letterforms distinct. For photographed documents, recognition reads a background-normalised copy so uneven phone lighting doesn't smear the letters together; that normalised copy is used for reading only and never appears in your output.

Recognition runs on the Tesseract engine, compiled to run inside your browser. The tool detects the document's script and language automatically, or you can choose from more than one hundred recognition languages, including right-to-left scripts such as Arabic and Hebrew. Pages that already contain selectable text are left unchanged. The summary lists the pages that were recognised and those that already contained text.

Check the result before download

After recognition, the tool reopens the generated PDF and checks it against the source. Page count must match; page dimensions and rotation must stay within 0.01 point; expected existing and newly recognised text must be searchable. It then renders every page and checks for visible changes. The result is offered only when all checks pass.

OCR is not added to a digitally signed PDF because changing the file would invalidate its signature. If every page already contains selectable text, no new copy is created and the source file can be used as it is.

Prepare a searchable archive

A single pass accepts up to 80 pages to limit memory use in the browser. Process longer archives in separate batches. Recognition runs on the device after the required language pack has downloaded; the selected document is not sent to a processing service.

Searchable PDFs can be indexed by an operating system, document-management system or desktop-search application. Before archiving a result, check names, dates, reference numbers and other critical text against the page image.

Processing on your device

The selected PDF, its pages and recognised text are processed in the browser. They are not sent to a Rebuild processing service, and the finished searchable PDF is saved through the browser as a new file.

Questions
When does a PDF need OCR?

Use OCR when text on a scanned page cannot be selected or found. It adds an invisible text layer for search, copying and indexing.

Which languages are supported?

More than 100 language models in the current Tesseract registry are available, including Turkish, Arabic, Indic scripts, Chinese, Japanese, Korean and historic variants.

What changes on the visible page?

OCR keeps the original page image and adds invisible searchable Unicode text. It does not rebuild the visible page content.

Where is OCR processed?

OCR runs in this browser. The selected recognition pack is downloaded and cached, while document pages are not sent to a processing server.