
Sole product designer on a tool that turns paralegal corrections into machine learning training data. Pre-trial discovery buries legal teams under thousands of scanned pages, and OCR misreads irregular scans badly enough that some of every page has to be checked by hand. When we trialled Tesseract against the scanner software the paralegals already had, they decided where that line sat: anything read below 86% confidence, they wanted to see. I designed and built the correction interface and the labeling schema. Our engineer owned the retraining pipeline.
Paralegals preparing degraded medical records for transcription had no signal about which pages would need manual work. Anything smudged or dirty was segregated and retyped whole, so a clean page and a ruined one cost the same read.
Corrections were kept as Word files. A correction held the retyped text and nothing else: not where on the page it sat, not what kind of error it was, not which field it belonged to. The model that produced the error never saw the fix.
The existing process was also mandatory and ours was experimental. Whatever came next had to fit inside a workflow we had no authority to change.
I shadowed three paralegals at an auxiliary office while they loaded scanners and set up OCR runs, digitizing badly degraded medical records. At the end of each day I interviewed them, and together we built a low-fidelity journey map to establish a baseline for human performance in the work.
We worked at the level of specific activities, loading the scanner, configuring the OCR tool, and tagged each one with an emotion, a thought, and the technologies a person was touching in that moment.
Then I gave each paralegal two votes to place on the moment of highest emotional load, and asked them to say why. Four went to manual review and two to waiting for the scan.

The interface is one station in a loop. Tesseract reads a scanned page and scores every passage it extracts; anything under the threshold the paralegals set is drawn on the page for review; the reviewer corrects the string and names the kind of error; and the correction goes back to the pipeline keyed by the index of the extraction it came from. The loop is the product. The screens below are the two places a person meets it.
The page, with its doubts drawn on it
When we trialled Tesseract against the scanner software they already had, the paralegals decided for themselves where the line sat: anything the OCR read below 86% confidence, they wanted to see and correct by hand. So the tool draws that set on the page. The idea came out of print design, where a proof is marked up in the margin of the thing itself rather than in a list somewhere else. The rest of the page is left alone, which is the whole change.
The panel is the notepad
Watching them transcribe, the same movement kept repeating: down to a notepad of lines and keywords, back up to the terminal, down again. They could type without looking at the keys, but not without looking away from the screen. The panel is that notepad moved onto the glass, one row for each doubtful passage. It collapses when it is not wanted, and its background stays transparent so the page underneath is never fully hidden by the notes about it.
The zoom had to be ours
Fine print on a degraded scan cannot be read at page size, and browser zoom pulled the outlines off the words they belonged to. So magnification is a control inside the tool, and the overlay scales with the document rather than against it.
Four words, and one of them is Other
The first round of corrections came back and told us that the corrected string alone was not enough to train on. Our ML architect insisted the list of error types be enumerable, so the paralegals worked out the four with him: Misread, Incomplete, Bad format, Other. They became the first features the model was given. Misread and Bad format are different failures, one of the model reading and one of the document's shape, and separating them is what lets the pipeline treat them as different evidence. Other is there so the set can stay small and still be honest about what it does not cover.
A correction leaves the interface already shaped for the pipeline
Every correction leaves the interface as one object: the index of the extraction it belongs to, what Tesseract read, what the reviewer put there instead, and the error type they picked. Keying on the index is what makes it cheap. The retraining step already holds the OCR output, so an index is enough to put the correction back on the exact region it came from, and nobody has to translate between the tool and the pipeline.
We designed a workflow that trains an OCR model on messy legal documents, instrumented around the low-confidence zones on each page. Across 1,250 training documents, confidence in those zones rose from 86% to 93%. The cost it was weighed against was concrete, roughly 12 low-confidence fields per document at 30 seconds each and a $35 hourly rate.
The most useful finding arrived after the workflow succeeded. Nobody had established ownership. Once it was demonstrably running, the questions surfaced: who controls the corrections, who may use the output in production, who signs off on retraining. No shared credentials existed, legal or technical, to answer them, and the project stalled there. Validating the model was half the job. Validating the relationships around it was the half nobody had scoped.
Confidence on low-confidence zones rose from 86% to 93% across 1,250 training documents. That is a confidence gain, not a verified accuracy gain: we ran no held-out comparison, so the claim is that the model got more certain, not that it got more correct. The cost it was weighed against (roughly 12 low-confidence fields per document at 30 seconds each, against a $35 hourly rate) is an estimate from observed sessions, not a time study.