Plain rule: OCR is a derived search and transcription layer. It can help locate a page or phrase; it does not replace the source image, PDF, filing, email, or other original record used to support a claim.
Source object first
The pipeline begins with the source object as received or acquired. Where the archive has durable file-level provenance, the original bytes, archive identifier, source path, and available integrity metadata are preserved separately from any text generated later. A re-render or re-OCR can change derived text without changing what the underlying source says.
Text extraction order
1. Native text
If a document already contains a usable text layer, native extraction is preferred. It avoids inventing OCR errors where machine-readable text already exists.
2. Page rendering
Image-only or degraded pages may be rendered to page images for machine reading. Page order and source-page boundaries remain part of the review context.
3. OCR derivative
Optical character recognition produces searchable text tied back to the source page. That text is treated as a derivative, not a certified transcript.
Normalization is intentionally limited
Whitespace and obvious extraction artifacts may be normalized for indexing, but names, numbers, dates, account references, dollar figures, tail numbers, addresses, and other evidentiary details are not silently “fixed” because a correction that looks obvious can still change meaning. When exact wording matters, the page image or original text layer controls.
Common OCR failure modes
- Character swaps: O/0, I/l/1, S/5, B/8, punctuation, and broken ligatures can change names, dates, identifiers, or amounts.
- Layout errors: multi-column pages, tables, marginal notes, stamps, headers, and footers can be read out of order.
- Image quality: skew, blur, bleed-through, low contrast, photocopy generations, handwriting, and compression can reduce recognition quality.
- Redactions: a black bar or missing region is a source condition. OCR does not authorize guessing what is underneath.
- False absence: a failed search hit does not prove a term is absent from the source; the OCR may simply have missed it.
How OCR hits become evidence checks
OCR is useful for discovery: finding candidate pages, clustering repeated terms, locating date ranges, and narrowing a large corpus. Before a material claim is published, the relevant source page should be opened and read in context. High-consequence details—especially identities, money, dates, legal language, and quoted wording—should be checked against the original record rather than accepted from an OCR hit alone.
Versioning and corrections
Derived text may be regenerated when a better scan, text layer, extraction method, or OCR pass becomes available. A new derivative should not be presented as if the underlying historical document changed. Corrections to transcription or indexing should preserve the distinction between source correction and derived-text correction.
Search and indexing boundary
Searchable OCR helps readers find records, but raw OCR dumps and thin machine-generated directory pages are not substitutes for source-specific editorial analysis. Indexable public methodology and investigation pages provide context, limits, citations, and verification paths; machine-derived text remains subordinate to the underlying record.
What this can show
- How a scanned or image-only record became searchable.
- Which text is native and which text is machine-derived.
- Why an OCR result should be verified against the source page before relying on exact wording.
- Why OCR can produce both false positives and false negatives.
What this does not prove
- An OCR hit does not prove the recognized spelling, number, date, amount, or quotation is exact.
- A missing OCR hit does not prove the term or fact is absent from the original record.
- OCR cannot reveal text hidden by a true redaction or reconstruct a page that was never provided.
- Digitization does not convert a source into a finding; interpretation still requires context and corroboration.
Related method guides
Review date
Method reviewed October 2, 2026.
This page describes the public evidence-handling rule for searchable text. If a specific record exposes a different extraction or provenance condition, the record-specific source note controls.