A records office: ring binders stuffed with paper stacked in rows across a long desk in the foreground, several of them split open under the weight, with two people working at benches further down the room behind them.

Training an ML Model to Improve OCR

A human-in-the-loop correction interface for legal discovery

Role Freelance
Type Freelance engagement · Sole designer, with one engineer on the pipeline
Evidence Project record: the shipped correction interface, its labeling schema, and the correction payload it emits.
Measurement Model confidence on low-confidence zones, 86 to 93%. A confidence gain, not a verified accuracy gain.

Overview

Sole product designer on a tool that turns paralegal corrections into machine learning training data. Pre-trial discovery buries legal teams under thousands of scanned pages, and OCR misreads irregular scans badly enough that some of every page has to be checked by hand. When we trialled Tesseract against the scanner software the paralegals already had, they decided where that line sat: anything read below 86% confidence, they wanted to see. I designed and built the correction interface and the labeling schema. Our engineer owned the retraining pipeline.

The Challenge

Paralegals preparing degraded medical records for transcription had no signal about which pages would need manual work. Anything smudged or dirty was segregated and retyped whole, so a clean page and a ruined one cost the same read.

Corrections were kept as Word files. A correction held the retyped text and nothing else: not where on the page it sat, not what kind of error it was, not which field it belonged to. The model that produced the error never saw the fix.

The existing process was also mandatory and ours was experimental. Whatever came next had to fit inside a workflow we had no authority to change.

Approach

I shadowed three paralegals at an auxiliary office while they loaded scanners and set up OCR runs, digitizing badly degraded medical records. At the end of each day I interviewed them, and together we built a low-fidelity journey map to establish a baseline for human performance in the work.

We worked at the level of specific activities, loading the scanner, configuring the OCR tool, and tagged each one with an emotion, a thought, and the technologies a person was touching in that moment.

Then I gave each paralegal two votes to place on the moment of highest emotional load, and asked them to say why. Four went to manual review and two to waiting for the scan.

A journey map of the paralegal transcription workflow: the stages of preparing scanned discovery for OCR, the actions taken at each, and where the reviewers said the work went wrong.
Drawn from shadowing three paralegals at an auxiliary office. Each activity carries an emotion, a thought, and the technologies in play, and each paralegal placed two votes on the moment of highest emotional load: four to manual review, two to waiting for the scan. The map shows seven dots, five on manual review, from a tallying slip in the session.

01 The ML training loop

The interface is one station in a loop. Tesseract reads a scanned page and scores every passage it extracts; anything under the threshold the paralegals set is drawn on the page for review; the reviewer corrects the string and names the kind of error; and the correction goes back to the pipeline keyed by the index of the extraction it came from. The loop is the product. The screens below are the two places a person meets it.

02 The training interface

The page, with its doubts drawn on it

When we trialled Tesseract against the scanner software they already had, the paralegals decided for themselves where the line sat: anything the OCR read below 86% confidence, they wanted to see and correct by hand. So the tool draws that set on the page. The idea came out of print design, where a proof is marked up in the margin of the thing itself rather than in a list somewhere else. The rest of the page is left alone, which is the whole change.

observed Capture of the shipped correction tool. The document is a filed court exhibit; personal fields are redacted in the capture.

The panel is the notepad

Watching them transcribe, the same movement kept repeating: down to a notepad of lines and keywords, back up to the terminal, down again. They could type without looking at the keys, but not without looking away from the screen. The panel is that notepad moved onto the glass, one row for each doubtful passage. It collapses when it is not wanted, and its background stays transparent so the page underneath is never fully hidden by the notes about it.

observed Capture of the shipped correction panel. The collapse and the transparent background are described from the working tool; neither is visible in a still.

The zoom had to be ours

Fine print on a degraded scan cannot be read at page size, and browser zoom pulled the outlines off the words they belonged to. So magnification is a control inside the tool, and the overlay scales with the document rather than against it.

observed Capture of the shipped correction tool at maximum zoom. The document is a filed court exhibit; personal fields are redacted in the capture.

03 The labeling schema

Four words, and one of them is Other

The first round of corrections came back and told us that the corrected string alone was not enough to train on. Our ML architect insisted the list of error types be enumerable, so the paralegals worked out the four with him: Misread, Incomplete, Bad format, Other. They became the first features the model was given. Misread and Bad format are different failures, one of the model reading and one of the document's shape, and separating them is what lets the pipeline treat them as different evidence. Other is there so the set can stay small and still be honest about what it does not cover.

observed Capture of the shipped correction panel. The four labels are the set as it shipped; no reviewer could add a fifth.

A correction leaves the interface already shaped for the pipeline

Every correction leaves the interface as one object: the index of the extraction it belongs to, what Tesseract read, what the reviewer put there instead, and the error type they picked. Keying on the index is what makes it cheap. The retraining step already holds the OCR output, so an index is enough to put the correction back on the exact region it came from, and nobody has to translate between the tool and the pipeline.

observed Capture of the shipped correction tool with its console open. The payload carries the four keys shown; the bounding box and confidence score stay on the OCR output side and are rejoined by index.

Outcome

We designed a workflow that trains an OCR model on messy legal documents, instrumented around the low-confidence zones on each page. Across 1,250 training documents, confidence in those zones rose from 86% to 93%. The cost it was weighed against was concrete, roughly 12 low-confidence fields per document at 30 seconds each and a $35 hourly rate.

The most useful finding arrived after the workflow succeeded. Nobody had established ownership. Once it was demonstrably running, the questions surfaced: who controls the corrections, who may use the output in production, who signs off on retraining. No shared credentials existed, legal or technical, to answer them, and the project stalled there. Validating the model was half the job. Validating the relationships around it was the half nobody had scoped.

Measurement note

Confidence on low-confidence zones rose from 86% to 93% across 1,250 training documents. That is a confidence gain, not a verified accuracy gain: we ran no held-out comparison, so the claim is that the model got more certain, not that it got more correct. The cost it was weighed against (roughly 12 low-confidence fields per document at 30 seconds each, against a $35 hourly rate) is an estimate from observed sessions, not a time study.

enesru