Messy lab reports in. One operator-ready table out.
Five real PDFs from five different labs — pathology, radiology, hematology, surgical margins, ED STAT panel. Each one written in a different layout, language register, and date format. Watch the parser land them all in a single normalized table.
5 real PDFs · click View PDF to inspectFictional patient data100% client-side · no API callsBuilt by Gerardo Espinosa
Step 1
Pick a sample report
01-oral-pathology-biopsy.pdf
24 KB
Westside Oral Pathology
Histopathology — oral biopsy · Houston, TX
LayoutPatient line is a one-liner with age as “47Y”. Diagnosis is buried at the bottom in a red box. Date written as “02-Apr-2026”.
LayoutRed STAT banner, mixed-alignment table. Flags as “CRITICAL UP / HIGH / LOW” and a separate critical-value notification footer. Date with time “18-May-2026 04:17”.
No reports ingested yet. Pick a sample above and click Ingest to watch the parser run and a normalized row land here.
Operator's manual
Each report carries an operator's note — the routing rule that decides whether the parse actually moves the case forward. Ingest a sample and the note for that case appears here.
How it works
What this demo does, how, and what's missing
What it does
→Ingests messy lab-report PDFs from 5 different labs — each with its own layout, field naming and date format.
→Normalizes every report into a single operator-ready row: patient, age/sex, date, lab, study type, finding, ICD-10, flag.
→Expands each row to show raw-extracted text side-by-side with the clean JSON record — so the parse is auditable, not a black box.
→Surfaces an operator's note per case — what the lab parses is only useful if the routing rule around it gets actioned.
How it does it
→All 5 PDFs are real files in /public/lab-samples/. Click View PDF to inspect the messy source.
→Each sample ships with both the raw extracted blob and the normalized record — pre-mapped by claude-sonnet-4-5 and shipped as static JSON. No API at runtime.
→The parsing console is a sequenced state machine that mirrors a real OCR → extract → normalize → flag pipeline; transitions are pure local state.
→Fully client-side. No backend, no keys, no network calls — safe to ship as a static portfolio page.
What a real impl would need
→Real OCR on scanned PDFs (Textract, GCV, or Tesseract) before Claude — most lab reports in the wild are flat raster scans, not text-embedded PDFs.
→Per-lab extraction templates + a fallback Claude pass with structured-output schema, plus a confidence score per field for the human-review queue.
→ICD-10 / SNOMED dictionary lookup with a curator step — the model proposes a code, a human accepts or overrides, and the pair becomes training data.
→Critical-value routing — pager/Slack/SMS for flag = critical, with acknowledgement tracking and an SLA timer per case.
→PHI handling: at-rest encryption, per-record audit log, retention policy, and a HIPAA/LFPDPPP review of the prompt + storage path.
→Eval set per lab template — when a new layout drifts, the regex / prompt regression shows up before patients do.
The pattern transfers to any document-heavy operator workflow — invoices, lab reports, intake forms, contracts. The value is in the routing rules around the parse, not the parse alone.