Rebuild guide

How to edit text in a scanned document without losing the original

A scan usually contains pixels, not editable words. The safest route is to identify what the file already contains, preserve the source, then choose OCR, reconstruction or manual editing for the result you actually need.

First, identify what kind of document you have

A digital PDF contains text and drawing objects created by software. You can usually select individual words, search the page and zoom without the letters becoming blocky. An image-only scan is closer to a photograph stored inside a PDF: it may look correct, but the words are not separate objects.

A searchable scan sits between the two. OCR software has added an invisible text layer over the page image, so search and copy may work even though the visible wording still cannot be edited cleanly. Testing selection, search and zoom before choosing a tool prevents unnecessary conversion and quality loss.

  • Digital PDF: selectable text and reusable document objects.
  • Searchable scan: page image plus an OCR text layer.
  • Image-only scan: pixels with no machine-readable text.

OCR is recognition, not complete reconstruction

OCR predicts which characters appear in an image. It does not automatically recover the original fonts, columns, table relationships, form fields, reading order or editable layout. Exporting OCR text into a blank word-processing file may be useful for quotation or search, but it is rarely a faithful replacement for the source document.

Document reconstruction goes further. It combines recognised text with layout analysis, page geometry and validation to create digital objects in the right places. A trustworthy system also keeps uncertain regions visible for review instead of silently inventing an answer.

Choose the result before choosing the workflow

If you only need to find words, a searchable PDF may be enough. If you need to correct one visible sentence, a PDF editor that can work with existing source text is a better fit. If the whole page is a scan and must become genuinely editable, it needs OCR plus layout reconstruction and a careful comparison against the original.

Forms, tables, handwriting, musical notation, maps and historical print need specialist handling. For those documents, preserving the complete page is safer than returning a confident-looking but incomplete digital version.

  • Search or copy: add a reviewed OCR layer.
  • Correct existing digital text: edit the original PDF objects where possible.
  • Recover an image-only page: reconstruct and validate the layout.
  • Change sensitive content permanently: use a real redaction workflow, not a white rectangle.

Use a source-preserving editing process

Keep the original file unchanged. Work on a copy, make the smallest necessary edits, and export to a new filename. If the editor must flatten a changed page, retain the untouched source separately and make sure hidden searchable text does not preserve wording that was meant to be removed.

For confidential material, check whether the tool processes the document locally or uploads it. Browser-local processing can keep ordinary file operations on the device, but a web page may still load analytics or advertising resources. Read the service's privacy explanation rather than assuming that every online editor handles files in the same way.

Review the exported document, not just the editor preview

Open the exported file in at least one independent PDF reader. Compare names, numbers, dates, punctuation, paragraph order, page size and line breaks against the source. Search for important terms, copy a sample paragraph, and zoom into tables or signatures to look for clipping and image degradation.

A visually convincing page can still have an incorrect text layer or reading order. For legal, financial, archival or accessibility work, the review should cover both appearance and machine-readable structure.

Where Rebuild stands today

Rebuild's public website currently offers working browser-local PDF and image utilities. The wider reconstruction system is being developed to preserve the source, recognise text, rebuild supported structure, expose confidence and let a person review the result.

Public web reconstruction and scanned-text editing are not available yet, so this guide does not send you to a pretend upload box. The iOS product is in beta development, and the website will publish an editing route only after its import, review and export paths work end to end.

Continue with Rebuild

Useful next steps