Graphiant Portal

We turned the pile of warnings an outage set off into one rewindable timeline per site. About 20 warnings became one, and engineers found the cause in under 30 seconds.

One feature, start to finish One warning timeline per site

  • Research found the real problem
  • Design designed the answer
  • Measurement measured what happened
Read the full study →

Context

Why did this work matter?

  1. C
  2. A
  3. R
  4. D
  5. I
  6. O

Graphiant's customers are network engineers watching tens of thousands of devices. When something broke, the old screen showed thousands of warnings and no clue which mattered, so engineers stopped trusting it.

The problem statement

The contract promises what we committed to. Today the portal how it falls short. That puts what is at risk at risk. We'll know it's fixed when the signal.

Filled in for monitoring

The contract promises monitoring that lets customers find, understand and resolve problems in their network. Today the portal fills the screen with thousands of warnings and no clue which one matters. That puts the engineers’ trust in the portal at risk. We'll know it's fixed when an engineer can name the circuit that is down in under 30 seconds.

Each product problem was written in this form before it went to design, worked backward from what the customer had been promised. Here the promise comes from Graphiant’s own proposal for monitoring, which asked for screens that let customers investigate and resolve any problem that came up.

Assumptions

What did we believe going in?

  1. C
  2. A
  3. R
  4. D
  5. I
  6. O

Before designing, we wrote down the hypothesis: one timeline per site would let engineers find an outage's cause in under 30 seconds. Writing it first meant nobody could move the goalposts.

The hypothesis template

We believe giving this feature to people doing this job will achieve this outcome. We'll know we're right when we see this signal.

Filled in for the warnings timeline

We believe giving one warning timeline per site to people doing triage when part of the network breaks will achieve finding the cause of an outage fast. We'll know we're right when we see an engineer find the cause in under 30 seconds.

Every design started as this sentence, with the last blank filled in before anyone opened a design file. Writing the signal down first meant nobody could quietly change what success meant later.

Why there is no persona in it

The template names a job, not a person, and that was on purpose. I had made personas for years, and by this point I no longer believed they did the work they were sold to do. A persona gives a made-up person an age, a photo and a backstory. None of that changed what a network engineer needed when part of the network went down. The job did: what they were trying to do, under what pressure, and how they would know they had done it. The portal had to work for whoever was on shift.

Demographic details also invite guessing. Once a pretend user has an age, a gender or a nationality, teams start designing for the stereotype instead of the evidence, a problem researchers have documented in how personas get used (Marsden and Haag, CHI 2016). If our research had shown that something about who the users were really changed the job, such as the language they worked in, it would have gone into the hypothesis as a finding. It never did. So the template left room only for what we could test: the feature, the job, the outcome and the signal.

Research

How was it tested?

  1. C
  2. A
  3. R
  4. D
  5. I
  6. O

We turned that number into questions a person can answer out loud: name the overloaded circuit, name the one that's down. Graphiant's sales engineers worked through a clickable prototype while we timed them.

Graphiant Validation Field Guide

Before the session

  • Hypothesis under test: We believe giving one warning timeline per site to people doing triage when part of the network breaks will achieve finding the cause of an outage fast. We'll know we're right when we see an engineer find the cause in under 30 seconds.
  • Pass mark, set before the first session: every participant shows they understand the screen in under 30 seconds. Thirty seconds is the failure line.
  • Participants: Graphiant sales engineers, standing in for customers. Note this on every result; where they and the live data disagree later, the live data wins.
  • Ask permission to record, and say how the recording will be used.

Introduction

  • Thank you for helping. This takes 30 to 45 minutes.
  • We're testing an early prototype, not you. Nothing here is finished, so criticism is the most useful thing you can give us.
  • Please think out loud as you go. I may not answer questions during the tasks; I'll come back to them at the end.

Warm-up

  • Tell me about the last time something broke on a customer's network. How did you find out?
  • What did you look at first, and what do you use today to find the cause?

Time to understanding (the clock starts when the screen appears; 30 seconds is the failure line)

  • What, if anything, is wrong on this network?
  • Which line would you deal with first, and why?
  • Record: seconds until they showed they understood, and whether their answer was right.

Find it (no hints, no clicking for them)

  • Show me how you would find out more about the problem you'd deal with first.
  • Record: reached the detail without help, yes or no.

Decide

  • What is happening on this line right now, and how long has it been happening?
  • Is it past its limits? How can you tell?
  • What would you do next?

Wrap-up

  • What did you expect to see that wasn't there? What would you take away?
  • How sure are you of your first answer, from 1 (guessing) to 5 (certain)?
Reading the map: one shaded patch, and the device in trouble.

Discoveries

What did we learn?

  1. C
  2. A
  3. R
  4. D
  5. I
  6. O

We set 30 seconds as the line where the design failed. The average was 8 seconds, nobody came close, and several sales engineers stopped the exercise once they got the idea.

Intervention

What changed because of it?

  1. C
  2. A
  3. R
  4. D
  5. I
  6. O

Warning history, live numbers and a picture of the devices sit in one place. A problem shows as one shaded patch on the map, and each site gets one timeline you can rewind and annotate.

Two views of the Cleveland HQ site, one overlapping the other. Behind, in the dark live view: Down 1h 5m ago, Edge 2 is offline, disrupting connectivity across all circuits, and Circuit 3 is down, impacting 256 applications, with Restart Edge 2, Reconnect Circuit 3 and Start Failover. In front, in the light rewound view: the timeline down the right edge, marked from 13:45 to 14:25 with warnings and errors, set back to 13:57, showing Warning 1h 38 mins ago, High Latency and Packet Loss detected on Edge 2, Edge may be approaching capacity, and a Return to Live button. Rebuilt for this site from late UI mockups, not a screenshot of the real portal, which shipped in 2024.
Live, Edge 2 is down. Rewound to 13:57, the warning that came first.

Outcome

What was the impact?

  1. C
  2. A
  3. R
  4. D
  5. I
  6. O

This one passed the prototype test, so it got built. In the live portal, about 20 warnings became one per site, and engineers found causes in under 30 seconds.

Also on this project

Read the full study →
  • 8 → 4 weeks

    Reusable setup templates halved the time to bring on a new customer's network.

  • +55%

    Task completion rose after common tasks came out of menus four or more clicks deep.

  • 3 designers

    I led three product designers; each design started as a hypothesis with its number first.

enesru