Training a Machine Learning Model to Read Messy Legal Papers

Paralegals chose to check only spots the OCR software was under 86% sure of, and as their fixes retrained it, its confidence there rose to 93% across 1,250 training documents.

One feature, start to finish The 86% confidence line paralegals set

  • Research found the real problem
  • Design designed the answer
  • Measurement partly measured

Partly measured: it shows how sure the model was, not whether it was right, because it was never tested on fresh pages.

Read the full study →

Context

Why did this work matter?

  1. C
  2. A
  3. R
  4. D
  5. I
  6. O

Before trial, legal teams get buried in scanned pages. OCR software turns a page picture into text but misreads badly worn medical records. Paralegals couldn't tell which pages needed checking, so anything smudged was retyped from scratch.

Assumptions

What did we believe going in?

  1. C
  2. A
  3. R
  4. D
  5. I
  6. O

I guessed that hand checking was the most stressful part, and that paralegals only needed to check spots the software was unsure about.

Research

How was it tested?

  1. C
  2. A
  3. R
  4. D
  5. I
  6. O

I shadowed three paralegals at a branch office, interviewed them each day, and drew a journey map of the work with them. Each got two votes for the most stressful moment.

A journey map of the paralegals' workflow: the steps of preparing scanned records for OCR, what they did at each, and where they said the work went wrong.
The journey map drawn with the paralegals; its seven dots come from a counting slip.

Discoveries

What did we learn?

  1. C
  2. A
  3. R
  4. D
  5. I
  6. O

Four of six votes went to hand checking. Trying Tesseract, a free OCR program, the paralegals set the line themselves: anything under 86% sure, they wanted to see.

Intervention

What changed because of it?

  1. C
  2. A
  3. R
  4. D
  5. I
  6. O

I designed and built a correction screen that outlines in red, right on the page, anything Tesseract rates under 86%. A side panel lists each doubtful spot as a field to fix, and each fix retrains the model.

The correction tool in a browser: a scanned emergency record on the left with every doubtful passage outlined in red, and a panel on the right listing each of those passages as an editable field beside a mistake-type menu.
Only spots under the 86% line are outlined in red.

Outcome

What was the impact?

  1. C
  2. A
  3. R
  4. D
  5. I
  6. O

Across 1,250 training documents, confidence on doubtful spots rose from 86% to 93%. It grew more sure, not proven more correct, and time saved was never measured. The project then stalled over who owned the corrections.

Also on this project

Read the full study →
  • A four-item list of mistake types worked out by the paralegals and our engineer.

  • A built-in zoom, because the browser's zoom pulled the outlines off their words.

enesru