Make a scanned PDF searchable
A searchable PDF can look unchanged while hidden OCR text drives search and copy. Test that layer before you depend on it.
Check whether the PDF is already searchable
Open the file in a PDF reader and search for a distinctive word. Select a sentence and paste it into a text field. If selection follows individual words and the pasted text matches the page, the PDF contains a text layer.
OCR can cover only part of a scan or sit out of line with the visible words. Test several pages, including numbers, punctuation and accented characters.
Understand the OCR layer
OCR predicts characters, words and reading order from page pixels. The page image remains visible while the predicted text is stored behind it, so copied text can contain errors that are not visible on the page.
Searchability does not make a PDF editable or accessible. Accessibility also depends on language, reading order, headings, table structure and descriptions; editing requires separate digital page objects.
Prepare the page for OCR
Rotate sideways pages, crop large empty borders and correct obvious perspective distortion. Heavy thresholding can darken print but also erase pencil marks, punctuation and fine table rules.
Choose the recognition language and account for mixed-language pages. Names, codes, abbreviations and historical type often need manual correction.
Check the hidden text layer
Search for several words, then copy paragraphs, numbers and table cells. Selection should follow the visible line, and extracted text should move through headings, body text, footnotes and columns in reading order.
For records that affect rights, money, identity or compliance, compare every critical value with the page image. Recognition scores do not replace that check.
- Names, dates, amounts and reference numbers match the page.
- Hyphenated line endings do not corrupt copied words.
- Multi-column text extracts in reading order.
- Invisible text is removed when content is permanently redacted.
- The PDF opens and searches in another reader.
Manage file size
OCR text adds little size compared with page images. Resolution, colour and image compression usually determine file size. Repeatedly converting a PDF to images can remove links, forms, tags and existing digital text.
Keep the source or archival copy separate from a smaller sharing copy, and make sure small text and important marks remain legible.
Save the source and searchable copies
Keep the untouched scan and save the searchable version under a different filename. Open the new file in another PDF reader and test search and selection. Create any smaller sharing copy from this version instead of repeating OCR.
For long-term records, note when the searchable copy was created, the OCR language used, which pages were checked and who corrected critical text. That record helps later readers distinguish the visible scan from machine-recognised wording.
Official sources
Sources checked on 14 August 2026 for the technical and procedural details cited in this guide.