A watercolour parcel map: lot boundaries in a street grid, washing from yellow through spring green to deep green

TerraQuotes

A patent-pending parcel-data estimator, and a pre-registered attempt to learn what a yard is worth before anyone visits it.

Shipped

Example estimate Parcel 4218-0031

Address 4218 Flora Place, St. Louis, MO 63110

Lot
13,916sq ft
Turf
9,240sq ft
Hardscape
1,845sq ft
Role Product & ML Engineer, Founder at Good Citizens
Type Own product · sole developer
Evidence Project record · repo read 23 Aug 2026
Measurement First-party telemetry, retained
Build Next.js · Python · Firestore · BigQuery · Mixpanel

Study Parameters

Design
Field-led, then instrumented
One homeowner’s question, a parcel index, three contractors.
Duration
17 months
First issue 2025-03-09; latest record 2026-08-21.
Registration
PR #934
Power-derived gate, registered 2026-07-14 before any training.
Primary metric
Win-probability AUC
Kill line 0.60 at around 400 outcomes. Derived, not chosen.
Mechanism
Parcel-derived pricing
Public land and building square footage. Patented.
Safety
Heuristic-first bound
A model may override only inside documented limits.

Overview

Founder and sole developer: I did the research, the product design, the Python parcel pipeline, and the application code. TerraQuotes tells you what landscaping your yard would cost, starting from your address. You type it in, pick your house off a list, and see what monthly upkeep runs. The number comes from the size of your lot with the buildings subtracted, which the county already publishes and gives away free. Three contractors use it today, and it is in production with money behind it.

The Challenge

I built a working demo in five months. The hard part was never showing a buyer a range of numbers. It was resisting the urge to build a massive DIY landscape design tool, and keeping laser focus on predicting which prices convert.

To learn that I need to know which estimates closed and which did not, and that lives in the landscaper’s own systems rather than mine: their invoices in QuickBooks, their deals in a CRM. A lost estimate only reaches me if someone inside their company turns on a setting I cannot see.

I had to proceed as if I would never obtain access to paid invoices, and built for that. The CRM as an adapter to QuickBooks exports, so an estimate could be compared to the invoice it became. Data modeling was never the problem; access to real volume was. And 978 hand-entered rows will not train a win-probability model: the gate is 75 wins and 225 losses before any training, and today it stands at 6 and 49.

Approach

I start from the assumption that I am wrong. Work is registered as a hypothesis with its null stated first, and it does not proceed until that null is rejected. I grade them by how bad it is if I am wrong and how much a user gains if I am right, and each is written down before the design, so the roadmap falls out of the ranking.

The riskiest hypothesis was that a free source of parcel sizes existed at all, and was good enough to price from. The whole product depended on it, and no national API served one. St. Louis published its own, land and building square footage both, and gave it away. The Python pipeline that turns that into a searchable address index is the part the patent claims, and it is what let the rest of the product start. The application is pending: the archive holds the specification, the claims, seven figures, and a mapping from each claim to the code that implements it. Pending is not granted, and no claim here rests on it having been.

Delivery runs from a backlog of specifications under an LLM-led protocol, so every feature is designed with its test suite first and verification is settled before code reaches production. Invalidation is ongoing: before each demo I build a field guide out of the design history file, and I run field research with every release.

Stop condition

The experiment will proceed only if all of these null hypotheses are rejected.

I tested three alternate hypotheses against a real price book when the first contractor, Quiet Village, came on. Their prices went into the model, and their customer records gave me the two questions I actually care about: does the estimate get closer to what jobs really sell for, and can a machine learn which quotes win.

  • Calculator trust

    A calculator cannot hold a sensible price while someone changes their mind about what they want.

  • Estimate price

    The number it lands on is not close enough for a landscaper to take the lead seriously.

  • Sq ft method

    A landscaper cannot price work by the square foot at all.

Assumption register
IDAssumptionCertaintyImportanceStatus
A-001Vercel's hosting will scale to all tenant traffic without us running our own serversHighHighValidatedN/A — validated in production
A-002A free geocoder is accurate enough for residential address lookup, so we never have to pay for oneHighHighValidatedSupplemented with county GIS parcel data as primary address source (CR-20251012)
A-003Routing geocoding through our own server removes the browser security and privacy problemsHighMediumValidatedN/A
A-004Public county parcel boundaries are authoritative enough to estimate service area fromHighCriticalValidatedAddress standardization handles heterogeneous schemas; YAML configs externalize field mappings
A-005Serviceable area can be derived reliably from public parcel dataHighCriticalValidated with caveatsvaluation and tier adjustments compensate for data gaps; area model semantics documented in CR-20260116
A-006Public property valuation is a meaningful proxy for landscaping price positioningMediumHighActive — monitoringPer-tenant pricing profiles can override; ML pipeline designed to learn from CRM outcomes
A-007Counting events on our own servers, in Firestore and BigQuery, gives enough insight without a tracking library in the browserHighHighValidatedBigQuery views provide analytical capability; Mixpanel receives PII-stripped UX events for real-time needs
A-008Labelling every event as either checking the build or testing the idea stays maintainable as the number of events growsMediumHighActive — enforcedCI enforcement via enforce-pr-contract.yml; event registry test validates purpose on all events
A-009Three storage tiers — Firestore for hot data, BigQuery for warm, blob storage for cold — is cost-effective at this scaleHighMediumValidated at current scaleCost monitoring in place; PostgreSQL migration plan documented in docs/tech-debt/MIGRATION_PLAN_OSS.md
A-010The people using this are homeowners on phones and crews on tablets, so mobile-first is the right defaultHighHighAssumed — reasonableDesktop progressive enhancement ensures both work; Chromatic tests both viewports
A-011Three service tiers (Curb Appeal, Full Lawn, Dream Lawn) match how residential landscaping is actually soldHighHighValidated operationallyPer-tenant tier labels configurable; telemetry tracks tier selection for Va
A-012Automated pricing is accurate enough for lead generation, though not for invoicingMediumCriticalActive — high priorityML pipeline (CR-20260209) designed to close the loop; CRM outcome detection (CR-20260306) will provide ground truth
A-013A one-time code by text or email is enough to gate estimate detail and capture a lead without driving people awayMediumHighActive — monitoringreCAPTCHA as secondary layer; rate limiting prevents abuse; verification events track completion rate
A-014One source of truth for design tokens, compiled per tenant, is enough to theme every tenantHighHighValidatedCompile-time is acceptable for current tenant count; revisit if dynamic tenant provisioning is needed
A-015Moving address search to the server removes the slowness caused by sending a 6MB search index to the phoneHighHighValidatedCase study documented in docs/case-studies/ADDRESS_INDEX_ARCHITECTURE.md
A-016A public mirror of lead data gives the sales team enough visibility without exposing personal detailsHighHighValidated with gapmirrorLeadToPublic Cloud Function trigger handles sync; gap indicates timing or error — needs investigation
A-017Naming the allowed host per tenant is enough to secure embedding without managing origins dynamicallyHighHighValidateddata-base-url attribute provides escape hatch for edge cases; TENANT_ORIGINS mapping (CR-20260311) handles CDN proxying
A-018Working out the tenant from the web address is enough for white-label hosting, without provisioning each oneHighHighValidated at current scaleAcceptable for current 3-tenant model; dynamic provisioning deferred to future EPIC
A-019Giving each tenant its own data path prevents one tenant reading another's leads, without row-level securityHighCriticalValidatedCross-tenant regression suite in place
A-020A baseline with a bounded override is more robust than putting the model firstHighCriticalActive — partially validatedHeuristic baseline always computed; ML is additive not required; CRM outcome data will enable validation
A-021Absorbing the typical bore cost, about $750, into irrigation base pricing produces cleaner estimates without underpricingMediumMediumActive — monitoring$750 based on typical one-bore scenarios; actual bore work priced correctly during installation
A-022Carrying CRM updates on events we already send avoids duplication and keeps the CRM interchangeableHighHighActive — partially validatedGeneric crm_* events designed for future providers; HubSpot namespace isolation prevents vendor lock-in
A-023Pulling modules out behind interfaces lets them be tested without touching the filesystem or the cloudHighHighValidatedRepeatable pattern documented for future module extraction
A-024Keeping each region's field mappings in validated configuration files is more maintainable than writing them in codeHighMediumValidatedJSON Schema export from Pydantic for IDE autocomplete is future improvement
A-025Feeding real outcomes back from the CRM will improve pricing accuracy by 15 to 25 percent within six monthsLowCriticalUnvalidated — infrastructure readyAll infrastructure in place; blocking dependency is CRM deal outcome detection (CR-20260306); measurement via BigQuery ml_hypothesis_metrics view
A-026Deals that were lost are as valuable as deals that were won for calibrating pricingMediumHighActive — manual capture shipped; native subscription pendingCR-20260306 epic closed 2026-03-23 (#348); manual capture path shipped July 2026; native loss-reason capture in build (#890)
A-027Partner sites will rewrite our embed script's address, so detecting where it loaded from is unreliableHighHighValidated (learned from incident)Hardcoded TENANT_ORIGINS mapping is deterministic and CDN-immune; data-base-url attribute for edge cases
A-028Dividing Phoenix, 500 square miles, by ZIP code is explainable to buyers and maintainable for exclusivity salesMediumHighActive — pre-productionConfig update process needed for ZIP changes; unmapped parcels route to parent maricopa region
A-029Creating the deal automatically improves close rate by at least 10 percent against keying it in by handLowHighUnvalidated — measurable via BigQueryBigQuery fact_lead_ml.deal_won rate by source enables measurement
A-030Feeding real outcomes back will close the gap between the estimate and what the job finally sold for, by 15 to 25 percent within six monthsLowCriticalUnvalidated — infrastructure readyMeasurement query defined in hypothesis registry; requires sufficient closed-deal volume first
A-031Aligning 2026 pricing to the official spreadsheet produces correct estimates for Quiet Village customersMediumCriticalActive — under validationyarn test:regression:quiet-village guards against drift; qvl2026PricingValidated telemetry tracks sign-off
A-032The stored area figure is already adjusted, so anything downstream must not adjust it againHighHighActive — fix in progressSemantic model documented; regression tests will validate fix

Every assumption, scored for certainty and importance. Load-bearing premises that were never tested are flagged in their column.

Registered gate

The model cannot start learning until enough finished jobs come back.

These numbers aren't random. I ran a simulation: I fed the model pairs of jobs, one won and one lost, and asked it to pick the winner. I tried different numbers of jobs, ran the simulation 30 times at each, and kept the smallest number of jobs where it picked the winner at least 80% of the time. Turns out that's 75 won jobs and 225 lost. Synthetic data can't tell me whether there is anything to find.

  • Losses recorded

    49/225

  • Wins recorded

    6/75

  • Wins we can train on

    0/75

Given one job that was won and one that was lost, the model put the winner first 47 times in 100 during July. A coin flip manages 50, and I said in advance I would give up below 60. Registered 2026-07-14, issue #1032.

01 Calculator trust

The square footage holds while you change your mind

The lot is measured once, and nothing on this menu re-opens it. You tick what you want, and every line says what it costs a month: mowing $2,221, fertilising $925. Cleanup is one choice out of four rather than four separate ones. Anything your plan does not cover shows a dash instead of a price you cannot buy. Then the totals, which are those same lines added up, kept apart so what you pay every month and what you pay once never blur together.

observed
Design input
GEM tiers = maintenance only
Action
CR-20260113-TIERS
Verification
Calculator returns separate totals; UI displays monthly maintenance + one-time enhancements + roll-in monthly
Outcome
REQ-20260111-006maintenancePackageSelected; enhancementToggled; rollInShown; appointmentRequested Va

02 Estimate price

The number, and the suite that keeps it honest

Full Lawn at $106,564, $16.30 per square foot, at 1012 Tallie Dr., Des Peres: the tier carousel over the resolved address, estimate pinned, breakdown one tap away

observed
Design input
Three-tier heuristic pricing
Action
CR-20250808-PRICING-001
Verification
engine.test.ts; formatters.test.ts; quantityResolvers.test.ts; pricingPolicy.test.ts; estimateTunables.test.ts; estimates.test.ts; pricingRegression.test.ts; api/estimate/route.test.ts; api/pricing/route.test.ts
Outcome
REQ-20250808-001estimate_generated; tier_select; service_toggle; roll_in_shown; tier_bundle_selected; pricing_engine_metrics Va

Where the leads land

The operator’s lead map clustering 15 estimates: cluster values of $44K and $74K, a count per marker, a filter control, and a header stating how many of the total are shown

observed
Design input
Lead map with filters
Action
CR-20251102-LEAD-CAPTURE-001
Verification
LeadMap.test.tsx; LeadMapFilters.test.tsx; LeadCaptureForm.test.tsx
Outcome
REQ-20251102-002lead_detail_copied; lead_filters_stage_changed; lead_filters_estimate_updated; lead_filters_date_updated; lead_filters_quick_action; lead_filters_reset Va

When I cannot price it, I say so

An address the index cannot resolve returns an explanation and a waitlist form instead of an error. Captured against Phoenix, which reaches this state through bug #1196; interim 504-guard shipped 4 Aug (#1201), prefix-bucketed rebuild underway (#1198–#1203)

observed
Design input
Tenant-scoped leads + public mirror
Action
CR-20251102-LEAD-CAPTURE-001
Verification
api/leads/route.test.ts; api/leads/[addressId]/convert/tenant-route.test.ts
Outcome
REQ-20251102-001lead_capture; lead_acquired Va

Every lead got a price

Thirty-eight people asked, and all thirty-eight got a number back the same minute. Nine booked a visit and four of those became work. The console shows the drop at every step rather than only the wins, and it says where its own figures stop: advertising spend is typed in by hand until an ad account is connected, and anything past won would need a different system again.

observed

03 Sq ft method

The lot, minus the buildings

Resolved address with a green chip showing 26,863 sq ft: serviceable area, lot geometry minus structures

observed
Design input
Multi-format shapefile ingestion
Action
CR-20250505-PARCEL-PIPELINE-001
Verification
test_shapefile_loader.py; test_csv_loader.py; test_region_config.py
Outcome
REQ-20250505-001Not measured

04 What the model waits on

The model is not allowed to run yet

The console's own readiness check, synced in May. Three gates have to pass before the model is allowed to retrain and one does, so it says Blocked and has never retrained. Its target of fifty recorded outcomes is the product's operational gate, read three months earlier than the 75 won and 225 lost registered above, and the accuracy gate it shows passing is estimate against final price, not the win or lose guess.

observed

A loss nobody wrote down teaches nothing

When the link to the contractor's own sales software fails, the estimate does not vanish. It falls into a short list somebody can settle by hand, and each row says what went wrong: one was refused by the calendar, the other never became a deal at all. The model can only learn from jobs that were recorded as lost, so this is where a loss gets recorded.

observed

Two connected, one gap that says why

Leads go out to the contractor's own sales software on their own, and won and lost results come back for training. The third card is not connected and states the reason rather than hiding it: nothing else in the stack reports what was actually invoiced or paid, so the funnel stops at won until that one is wired up.

observed

05 The shipped app

The lead map

Every lead that has entered the pipeline, placed on the parcel it came from. The counters above it are the whole funnel in three figures: estimates, pipeline value, and the share that reached the CRM.

observed Capture of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Browser chrome cropped; nothing else altered.

Clustered where the work is

Zoomed in, the clusters break apart into individual parcels. Density is the point: a route is worth driving when the jobs are near each other.

observed Capture of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Browser chrome cropped; nothing else altered.

Density, not pins

A heat layer answers a different question from a pin: not where is this lead, but where is the work concentrated.

observed Capture of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Browser chrome cropped; nothing else altered.

Filtered by stage

All, New, Contacted, Archived. The map is a view of the pipeline rather than a separate list of its own.

observed Capture of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Browser chrome cropped; nothing else altered.

Filtered by time

All time, today, last 7, 30 or 90 days, or a custom range. What arrived this week is a different question from what is on the books.

observed Capture of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Browser chrome cropped; nothing else altered.

One estimate, opened

The record behind a marker: the price band, the property type, the measured area, the address it was taken from, and the two buttons that close it out. Marking it won or lost is what feeds the model.

observed Capture of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Browser chrome cropped; nothing else altered.

Out to the books

An accepted estimate leaves the tool as an invoice. That handoff is the reason the revenue figures on the map can be trusted: they are the same numbers the business bills against.

observed Capture of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Browser chrome cropped; nothing else altered.

Lead to won

Twenty-seven leads in, one instant estimate, six consultations booked, two won. The gap between estimate and booking is the funnel; the tool states it rather than burying it.

observed Capture of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Browser chrome cropped; nothing else altered.

Where they came from

The same pipeline cut by source: direct, the services page, the lawn-care page, HomeAdvisor. A lead with no CRM link is shown as exactly that.

observed Capture of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Browser chrome cropped; nothing else altered.

The model is not allowed to run yet

Blocked, and the screen says why. Four gates, each with its own threshold, and the count of resolved losses well short of the one that matters.

observed Capture of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Browser chrome cropped; nothing else altered.

The gate that is holding it

49 of the outcomes it needs. The number is on screen because a model that trains on too little is worse than one that has not trained at all.

observed Capture of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Browser chrome cropped; nothing else altered.

Two connected, one gap that says why

GoHighLevel and a CRM are wired in; the field-service integration that would carry invoiced and collected back is not. The screen names the gap instead of implying coverage.

observed Capture of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Browser chrome cropped; nothing else altered.

The estimator, embedded

The same estimator runs on the contractor's own site as a tagged widget, so a lead arrives already attributed to the page that produced it.

observed Capture of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Browser chrome cropped; nothing else altered.

06 The estimator, moving

The stills show the estimator's states; this shows the path a homeowner actually takes through it. One address in, a priced package out, without a sales call in between.

observed Screen recording of the shipped application at quietvillagelandscaping.com/gem-concierge-quote-estimator, August 2026. Re-encoded for delivery; nothing else altered.
  • Address to monthly total

    Sixty-three seconds, no sales call. The parcel does the arithmetic the estimate is built on.

Outcome

H-156 was registered with its null stated alongside it, so the result had somewhere to land if the loop did nothing: “There is no meaningful improvement in estimate accuracy from the ML training loop.” The alternative was that feeding real outcomes back would close the gap between the estimate and what the job finally sold for by 15 to 25% within six months. If the loop ran and the gap did not move, the null stands and the pipeline is expensive plumbing.

Neither happened. The comparison the null and the alternative both depend on turned out to be arithmetic between two different units: the retired columns mixed a monthly figure with two one-time figures on 215 of 347 historical rows, with nothing to tell them apart. It was withdrawn before either could be scored. A null that cannot be tested is not the same as a null that holds, and the registry says which of the two this is.

What replaced it is not another guess at a number. The 15 to 25% was chosen, and the register calls it aspirational; the gate the model now answers to is derived. A power simulation registered on 2026-07-14 fixes the cohort before any training, at 75 won and 225 lost resolved outcomes, and states in advance the condition under which the model is not worth building at all: AUC below 0.60 at around four hundred. The measurement is a BigQuery view written before the rows exist rather than a query composed once the answer is visible. Today the gate stands at 49 losses against 225 and six wins against 75, with no usable won rows yet.

The machinery around it works. A model may move a price only inside its documented limits, and that path has been exercised in production. What is not settled is whether the losses arrive at all: the open action on the cohort review is to confirm that lost and abandoned deals are landing from both contractors. If they are not, the negative class is being dropped at the source, and nothing decided later recovers a loss nobody wrote down. That is the part I do not control, and it is the part the model is waiting on.

Measurement note

Figures are first-party telemetry from the product's own BigQuery views, read from the repo on 23 August 2026. The cohort gate (49 of 225 resolved losses, 6 of 75 wins) is a live count against a power simulation registered on 14 July 2026, before the rows existed.

The estimate-accuracy hypothesis H-156 was withdrawn, not answered: the comparison it depended on mixed a monthly figure with two one-time figures on 215 of 347 historical rows, with nothing to tell them apart. A null that cannot be tested is not a null that holds. No model has been trained, so no accuracy claim is made here at all. Three contractors in production is a count, not a retention figure.

The patent is an application, not a grant. Nothing in this record depends on it issuing.

Traceability

Sixty-three requirements. Each lists its acceptance criteria, the tests that verify it, the events that validate it in production, and any remaining gap. Nineteen ship with a named gap, printed here exactly as the CSV records them.

Traced requirements
63
Shipping with a named gap
19
Scored assumptions
32
Dated artifacts
166

TerraQuotes is raising a seed round. Every requirement above is traced from acceptance criteria written before the build, through the tests that verify it, to the events that validate it in production. Nineteen ship with a gap that is named rather than omitted, and the pricing model is held behind a threshold registered before the data existed. The design history file runs from 2025-03-04 to 2026-08-21. Diligence here does not start with a deck: ask for the packet and read the record, including the parts that are still open.

bob@goodcitizens.us
enesru