Training a Machine Learning Model to Read Messy Legal Papers
Paralegals chose to check only spots the OCR software was under 86% sure of, and as their fixes retrained it, its confidence there rose to 93% across 1,250 training documents.
One feature, start to finish The 86% confidence line paralegals set
- Research found the real problem
- Design designed the answer
- Measurement partly measured
Partly measured: it shows how sure the model was, not whether it was right, because it was never tested on fresh pages.
Read the full study →Context
Why did this work matter?
- C
- A
- R
- D
- I
- O
Before trial, legal teams get buried in scanned pages. OCR software turns a page picture into text but misreads badly worn medical records. Paralegals couldn't tell which pages needed checking, so anything smudged was retyped from scratch.
Assumptions
What did we believe going in?
- C
- A
- R
- D
- I
- O
I guessed that hand checking was the most stressful part, and that paralegals only needed to check spots the software was unsure about.
Research
How was it tested?
- C
- A
- R
- D
- I
- O
I shadowed three paralegals at a branch office, interviewed them each day, and drew a journey map of the work with them. Each got two votes for the most stressful moment.

Discoveries
What did we learn?
- C
- A
- R
- D
- I
- O
Four of six votes went to hand checking. Trying Tesseract, a free OCR program, the paralegals set the line themselves: anything under 86% sure, they wanted to see.
Intervention
What changed because of it?
- C
- A
- R
- D
- I
- O
I designed and built a correction screen that outlines in red, right on the page, anything Tesseract rates under 86%. A side panel lists each doubtful spot as a field to fix, and each fix retrains the model.

Outcome
What was the impact?
- C
- A
- R
- D
- I
- O
Across 1,250 training documents, confidence on doubtful spots rose from 86% to 93%. It grew more sure, not proven more correct, and time saved was never measured. The project then stalled over who owned the corrections.
Also on this project
Read the full study →A four-item list of mistake types worked out by the paralegals and our engineer.
A built-in zoom, because the browser's zoom pulled the outlines off their words.