← Back to portfolio
Case study

Taking a healthcare document-extraction pipeline from ~70% to ~100% accuracy

Company: Evercred, a healthcare credentialing platform (physician, enterprise and admin apps). The company shut down in Sep 2026.

My role: Senior Full-Stack & AI Engineer. The extraction pipeline already existed but wasn't reliable in production. I took over the technical side, fixed it and completed it into a working production version. I worked with a product owner who set priorities and a dedicated QA team.

Stack: Claude API (vision) · Python/Django · AWS S3 · PostgreSQL (Prisma, multi-schema) · Next.js 14 · Turborepo monorepo · Terraform + EKS

Architecture only. No proprietary code, client data or internal names.

The problem

Credentialing starts with paperwork: medical licenses, certifications and IDs, each in a different format depending on the state or issuing body. Someone has to copy license numbers, names and dates into a system accurately, because a wrong license number means a failed verification and a delayed credential.

For context on why this matters (industry figures, not Evercred's):

Evercred's pipeline was meant to remove the manual copying: upload a document, get structured fields back, confirm them.

How the pipeline works

Upload (web app)
   │
   ▼
S3 storage
   │
   ▼
Python/Django preprocessing: page splitting and image prep (not traditional OCR)
   │
   ▼
Claude API, vision-based extraction → structured fields (license number, name, dates)
   │
   ▼
Validation layer ── per-field format checks
   │                     │ fails?
   │                     ▼
   │               Cross-check against the PDF's embedded text layer
   ▼
PostgreSQL
   │
   ▼
Human review: the practitioner confirms or corrects the extracted data in the UI
  1. Intake. Documents are uploaded through the app and stored in S3.
  2. Preprocessing. A Django service splits pages and prepares images for the model.
  3. Extraction. Claude reads the document image directly and returns structured fields. There's no separate OCR step, so the model sees the layout the way a person would.
  4. Validation. Every extracted field is checked against the format rules for its type. This is where the accuracy fix lives (below).
  5. Human confirmation. The practitioner reviews and confirms the extracted data before it's treated as final. Nothing extracted by the model is accepted silently. (When a credential was shared with an enterprise, someone at that organization could also verify it. That was a separate flow outside this pipeline.)

The accuracy fix (the centerpiece)

Symptom: on some license-number formats, extraction was correct only about 70% of the time. For a credentialing product that's unusable: every wrong number is a verification that fails later, in front of a customer.

Fix: I added a validation step with a fallback source of truth.

  1. Validate each extracted license number against its expected format.
  2. If validation fails, don't trust the vision output. Check it against the PDF's embedded text layer, which is the actual text inside the PDF file, when the document has one.
  3. Use the confirmed value, then pass it to human review as usual.

Result: accuracy on those formats went from ~70% to ~100%.

Why it works: vision models are strong at layout and weak at being exactly right on long alphanumeric strings. Many PDFs already carry the exact characters in their text layer. Using the model for understanding and the text layer for exactness, and only when validation says something's wrong, gets the benefits of both without slowing down the documents that were already fine.

What I'd add for a HIPAA-grade deployment

These weren't part of the Evercred pipeline. They're the next layer I'd build for a product that handles patient health data, depending on its compliance requirements:

  • Audit trail: who uploaded, viewed, edited or confirmed each field, and when, including which model version produced the value.
  • Model access under a BAA (the vendor agreement HIPAA requires), for example Claude via AWS Bedrock or Anthropic's HIPAA offering.
  • No patient data in logs or traces, with redaction at the logging layer.
  • Measured accuracy as a CI gate: an evaluation suite that scores each field against labelled documents, so a prompt or model change can't quietly regress.