Source guide
Provenance, OCR and duplicate files
How TRP handles multiple copies, OCR text, mirrors, redactions and later versions of the same underlying record.
Key takeaway
One underlying document can have many representations. More copies do not automatically mean more evidence.
Document identity and representation are different
The same underlying record may exist as an official release, a court copy, an archive mirror, an OCR text layer, an endorsed copy, a redacted version or a later republication.
TRP aims to preserve that lineage so readers can tell whether they are looking at new evidence or another representation of evidence already seen.
OCR is fallible
Searchable text can misread names, dates, handwriting, columns and redaction marks. OCR is therefore a discovery and accessibility layer rather than a substitute for the source image.
When a quotation or exact wording matters, the source page controls.
Duplicates do not multiply corroboration
Exact hashes, normalized similarity and page-level comparison help identify duplicate families and later versions.
A repeated page receives no additional evidentiary weight simply because it appears in multiple releases or mirrors; genuinely new endorsements, annotations or pages are treated as deltas.
