Rebuild guide

How to make a scanned PDF searchable—and verify the result

A searchable PDF usually keeps the original page image and adds recognised text behind it. That can unlock search, copy, indexing and assistive workflows, but the hidden text still needs to be checked.

Check whether the PDF is already searchable

Open the file in a normal PDF reader and search for a distinctive word that appears on the page. Try selecting one sentence and copying it into a plain-text field. If selection follows individual words and the pasted text is accurate, the document already contains a text layer.

Some scans contain incomplete OCR: headings may search while body text does not, or the text layer may be offset from the visible words. Test more than the first page and include numbers, punctuation and accented characters.

What OCR adds to a scanned page

OCR analyses page pixels and predicts characters, words and sometimes reading order. In a searchable PDF, the original image remains visible while predicted text is stored invisibly in approximately the same locations. The visual page can therefore look perfect even when copied text contains errors.

Searchability is not the same as editability or accessibility. A useful accessible document also needs correct reading order, language, headings, table structure and descriptions where appropriate. A genuinely editable document needs digital layout objects rather than only an invisible overlay.

Improve the source before recognition

Recognition works best with sharp, upright text, adequate resolution and even contrast. Rotate sideways pages, crop large irrelevant borders and correct obvious perspective distortion. Treat heavy thresholding carefully: it can make dark print clearer while erasing light pencil, punctuation or fine table rules.

Select the correct recognition language and keep mixed-language material in mind. Names, codes, uncommon abbreviations and historical type often need manual correction even when ordinary prose performs well.

Validate the hidden text layer

Search for several words on every document type, then copy representative paragraphs, numbers and table cells. Check that selecting text follows the visible line rather than jumping across columns. Screen-reader or text-extraction order should move through headings, body text, footnotes and columns in a sensible sequence.

For records that affect rights, money, identity or compliance, compare every critical value against the image. OCR confidence is a useful review signal, not proof that the content is correct.

  • Names, dates, amounts and reference numbers match the page.
  • Hyphenated line endings do not corrupt copied words.
  • Multi-column text extracts in reading order.
  • Invisible text is removed when content is permanently redacted.
  • The result opens and searches in an independent PDF reader.

Keep file size and visual quality under control

OCR text usually adds little size compared with page images. File size is driven mainly by image resolution, colour and compression. Reduce size only as far as the smallest important text remains readable; repeatedly converting a PDF to images can remove links, forms, tags and existing digital text.

Keep an archival or source-quality copy separate from a smaller sharing copy. A compact PDF is convenient, but it should not become the only surviving version if compression has removed evidence.

Rebuild's current searchable-document path

The public Rebuild website does not yet offer OCR or a make-searchable upload route. Its current tools process supported PDF and image operations locally in the browser. The planned Rebuild workflow goes beyond an OCR overlay by preserving the source, reconstructing supported text and structure, recording confidence and requiring review before export.

Until that route is complete, this page remains an honest guide rather than a placeholder tool. You can use the available browser utilities for page organisation and conversion, or follow the product status page for platform availability.

Continue with Rebuild

Useful next steps