Esc

↑↓ move↵ openIndex · Pagefind
Technology

OCR (Optical Character Recognition)

Definition

Technology that reads text from scanned documents, PDFs and photos and turns it into data software can use.

OCR converts images of text, such as a scanned invoice, a photographed delivery ticket or a PDF lien waiver, into machine-readable text. Modern document capture goes further, combining OCR with models that find specific fields (vendor, invoice number, date, amount, PO number, cost code) and return structured data.

Why it matters

Construction runs on paper and PDFs: invoices, T&M tickets, waivers, delivery tickets, certified payroll. Typing them in is slow and error-prone. Good OCR turns a stack of PDFs into draft records, so people review and approve instead of keying every field.

Worked example

Illustrative: Volt Electric emails INV-4471. OCR extracts vendor, invoice number, date, $48,200 and a reference to PO-118. The system matches it to the PO and sees only $42,050 left on the commitment, so the invoice is $6,150 over. Instead of posting it, it’s routed to the PM with the variance highlighted.

Common mistakes

  • Trusting extracted amounts without a review step, especially on handwritten tickets.
  • Measuring OCR accuracy per document instead of per field that matters.
  • Capturing data but not matching it to the PO, subcontract or cost code it belongs to.

How os.construction handles it

We’re building document capture that turns invoices and waivers into draft records matched to commitments, with a person approving before anything posts. It’s in development with founding contractors.