Organise PDF

Compare two PDFs

Compare PDFs across text, layout and appearance, with optional OCR for scans.

Working…

Compare two PDFs

Match corresponding pages, separate text, layout and visual differences, review uncertain scan evidence and export a report linked to both source files.

Confirm the page pairing

The first selected PDF is the original and the second is the revision. Extracted text and compact visual signatures help match corresponding pages, while unmatched pages are listed as additions or removals.

Each source page is accounted for once. When a likely match is ambiguous, the page pair is presented for confirmation so a missing page does not automatically shift the comparison of every page that follows.

Pairing selects the pages to compare; alignment places their crop boxes and rotations in a shared coordinate system. A changed page size or rotation remains a layout finding, and the complete page map is available for review.

Separate text, layout and visual changes

For a digital PDF, the embedded text layer supplies words and positions. Added, removed and replaced runs are content findings; unchanged words that move, and differences in page geometry, are layout findings.

Links, form fields and annotations are compared as structural objects. A separate page render can identify changes to artwork, colour, rules and images that are not represented in extracted text.

A visual region that supports an existing text or layout finding is attached to that finding instead of being listed repeatedly. Remaining regions stay visual, with fixed tolerance for rasterisation and anti-aliasing noise and a confidence label for the classification.

Review scanned pages with optional OCR

A scanned page without embedded text is compared visually first. When either page in a matched pair lacks words, the results offer an explicit local OCR pass for one selected language rather than starting recognition automatically.

Each OCR page targets a 300 DPI render before the 2,000-pixel long-edge and 2,000,000-pixel safety caps. Recognition runs sequentially on no more than ten page pairs, and a model or recognition failure leaves those pages labelled visual-only.

OCR evidence remains separate from native PDF text and page-matching evidence. Possible wording changes use estimated regions, and equal recognised wording still requires review because names, numbers, dates, handwriting, tables and faint print can be misread.

The report records the language, model and method hashes, per-page evidence hashes and incomplete pages. Every review pair must be opened and acknowledged before the PDF or JSON report can be exported.

Export reports tied to both source files

Comparison creates a separate report and does not overwrite either PDF. It records each source filename, byte size, page count and cryptographic hash, plus the comparison and schema versions, confirmed page map, settings and limitations.

Each finding includes its page pair, category, evidence and bounded before-and-after text when native or recognised words are available. The report does not validate a digital signature, legal effect, authorship or source authenticity.

A job accepts exactly two unlocked PDFs, up to 100 pages each and 40 MB together. The browser reads and renders one bounded pair at a time, and files, extracted wording and hashes are not sent to a comparison server.

Export the PDF report for human review and the structured JSON report for a retained audit trail or later automated checks. Keep both source PDFs with the report so its hashes and page references can be resolved.

Questions
What happens to the source PDFs?

Comparison creates a separate report and leaves both source PDFs unchanged.

How are scanned PDFs compared?

Scans are compared visually first. Optional OCR covers up to 10 matched page pairs where either side lacks text; recognised names, numbers, dates, handwriting, tables and faint text require visual review.

Where is comparison processed?

Comparison and optional OCR run in this browser. The selected language model is fetched from Rebuild, while the PDFs and recognised text remain on the device.