
I led three product designers on Graphiant's network management portal. Graphiant sells enterprise networking to large companies, and their customers' network engineers use this portal to set up and watch over tens of thousands of pieces of hardware.
I had a second job running alongside the shipping work: change how the team decided things. Before, designs were argued. After, they were tested against a number written down in advance.
The portal already existed, and it did not work. Setting up a network took eight weeks. The live map of the network rendered empty. The bulk upload had never worked. Common tasks sat four or more clicks deep in nested menus.
The engineers who used it stopped trusting it. That kind of failure never reaches a bug tracker, so another round of redesign by opinion would not have found it either. The team needed a way to tell, before and after each change, whether the change had helped.
Every problem went through the same four steps, inside the sprint schedule the product managers already kept.
One: write the bet down. Before anyone opened a design file, we wrote what we expected to happen and the number that would settle it. How many people finish the task. How long it takes to find the cause of an outage. Writing it first meant success could not be redefined later. When the team disagreed, we wrote both versions down and tested both.
Two: turn that number into a question a person can answer out loud. Name the circuit that is overloaded. Name the one that is down. Show me where you would click for detail. Designers, product managers and engineers wrote these together and kept them in a shared field guide.
Three: design against it. Reusable templates replaced the broken bulk upload. The network and its traffic were drawn live instead of left blank. Thousands of stacked alerts collapsed into one timeline per site you could rewind. Tables were rebuilt so an engineer could act on many devices at once.
Four: test before building. Graphiant's sales engineers troubleshoot the product for real customers daily, which made them the closest thing to a customer we could get in a room. They worked the tasks while we scored them against what we had written.
Only designs that passed got built. Those same task definitions became the analytics we shipped, so the measuring carried on in production.
Triage as a picture, not a pile
The registered hypothesis: unify alert history, streaming metrics, and the device diagram, and anomalies stop hiding: glyphs, motion, and stroke styles carry connection health; a timeline adds rewind and annotation. Validated with 30-second glance tests, recall and interpret the state of the network, then by time-to-root-cause telemetry in production. It closed at a 20:1 signal-to-noise ratio with root cause identified in under 30 seconds. Rebuilt in PrimeVue in 2025; the original shipped in 2024.
Tables rebuilt for fleets
The original tables were designed for one device at a time; no bulk actions, no sorting, no way to work a fleet. Rebuilt with multi-select, filtering, pagination, resizable columns, and sub-interface inspection. Validated in moderated SE task runs against the field guide's table prompts, with the timing confirmed in Mixpanel: table tasks completed 40% faster. Rebuilt in PrimeVue in 2025; the original shipped in 2024.
Bulk actions replace CSV uploads
Multi-select and a bulk Actions menu did what the CSV importer never could: change many devices at once, inside the UI, with a confirmation you can read. Validated first as SE prototype tasks, then by production adoption telemetry: use of the bulk-configuration tools tripled once they existed. Rebuilt in PrimeVue in 2025; the original shipped in 2024.
Every physical port, mapped
Dozens of virtual interfaces, and no way to see a physical port's status in a table-only layout. Mapping every port in the UI: each button opening its context-specific form and sub-interface table. Validated in moderated triage tasks with sales engineers, with QA backstopping the edge cases; Mixpanel interaction counts confirmed it in production: triage twice as fast, sub-interface table interactions up fivefold. Rebuilt in PrimeVue in 2025; the original shipped in 2024.
Four recordings of the design system working the field guide's own tasks: the same prompts sales engineers were tested against, from 30-second topology glance tests to a bulk configuration push. All of them are reconstructions: the interface rebuilt in PrimeVue for Vue.js, the component library the team's PM selected, because the shipped portal lives behind customer logins. The stills show where each screen arrives; these show how it gets there.
Finding the anomaly
The field guide's triage tasks, verbatim: name the overloaded circuit, name the circuit that is down, where would you click for detail. Validation method: 30-second glance tests; sales engineers viewed the topology briefly, then recalled and interpreted the network's state. The registered criterion was root cause in under 30 seconds; this is the path that meets it, one shaded segment instead of a thousand stacked alerts.
Editing four devices at once
The flow the CSV importer was supposed to be: select, edit once, push to all, read the confirmation. Validation method: moderated SE prototype tasks pre-build, then Mixpanel adoption telemetry post-ship; bulk-configuration tool usage tripled.
Region to site to device
The navigation restructure in one pass (region, site, device), with the incident carried along instead of re-hunted at each level. Validation method: moderated task runs against the field guide's navigation prompts, then Mixpanel's reset and completion events in production; task completion rose 55%, home resets fell 25%.
Configuration as objects
The template concept underneath the onboarding numbers: circuits, devices, and interfaces as reusable objects rather than rows in a spreadsheet. Validation method: alternative template hypotheses tested as SE prototype tasks, with QA backstopping edge cases; Mixpanel reuse logs carried the winner into production; setup time halved, 100,000+ devices on shared templates. Recorded from the light-mode prototype of the same system.
Two figures carry this engagement. Enterprise network onboarding time fell by half, from eight weeks to four, using centralized configuration templates; Mixpanel logs confirmed reuse of those configuration objects across more than 100,000 devices. And alert aggregation reached a 20:1 signal-to-noise ratio, collapsing thousands of events into one actionable alert per site, with root cause identification under 30 seconds.
The rest of the register closed too, more quietly: task completion up 55% and home resets down 25% after the navigation restructure, table tasks 40% faster with bulk configuration use tripled, and port mapping making interface triage twice as fast. Each of those answers a criterion that was registered before its design sprint began. The instrumentation existed because the hypotheses demanded it, not the other way around.
the product figures are first-party telemetry against pre-registered criteria, with baselines and measurement windows recorded in the project telemetry rather than reconstructed here. Sales-engineer sessions were moderated tests against the field guide's written task prompts, and a proxy is not a customer; where SE behavior and Mixpanel production data disagreed, production data won.
The practice did not outlast the engagement. Hypothesis registration and the field guide were introduced and used, the figures above are what they produced, and then they were declined quietly, by not being used. The framework asked product managers to learn something before a release date they had already committed to, and it offered no cheaper way to keep that commitment.
The need underneath was real, and they named it: a way to track what their product decisions actually did, without a research budget. I proposed one place where a decision, the evidence behind it and the version it shipped in all sat together. The idea landed. The habit did not.
What the company wanted from measurement was a figure that was safe to repeat outside the building. That is a different instrument from one built to change your mind. Measurement in client work gets attached to how a team is performing; the thing worth measuring is how the product is performing. The figures above are the second kind. The practice that kept producing them was the first kind's problem, and it lost.
This account of why it was not adopted is mine, written afterwards, with no artifact in this archive to check it against.