The Pour-Over Protocol

A pre-registered A–B–A self-experiment in drinking less coffee without giving anything up

Concluded

Role Lead Product Designer at World Wide Technology
Type Self-experiment, n = 1: participant and experimenter are the same person
Evidence Diary + coded observation + cross-validation log; 19-file packet, downloadable
Measurement Diary ran the length of the study; observation and cross-validation ran during intervention Weeks 3–6
Clearance Physician continuation clearance at Week 4 (2024-01-27), in the packet

Study Parameters

Design
A–B–A, n = 1
2 weeks baseline · 4 intervention · 2 washout
Duration
8 weeks
2024-01-06 → 2024-02-21 on record
Registration
PRE-001
2024-01-05T09:00:00Z
Primary metric
cups_per_week
Baseline measured by 14-day observation diary
Mechanism
Desirable difficulty
Manageable friction that slows short-term performance but strengthens habit formation
Safety
Physician-gated
Mandatory review at systolic >150 or diastolic >95

Overview

Sole author and sole subject of an n=1 self-experiment: I designed the protocol, built the logging instrument, and was the participant. My doctor put me on blood-pressure medication and told me to cut back on coffee. The reason she gave was work: the stress of the job I was doing at the time, and what it was doing to my blood pressure. Coffee was the lever she could name. She did not say by how much. Heck, I didn't even know how much I was really drinking. All I knew is that I guessed around 3–4 cups every day. But before changing anything, I captured a true baseline by writing down every cup I drank for two weeks. The answer was closer to 32 a week; 4.57/day, 40% above my own estimate. "Way more than I thought," the decision log says. I wanted to cut back, but I didn't want to replace coffee with something worse. So this study pulled a different lever: not a limit on cups I could have, but a harder way to make each one.

The Challenge

Before the first pour-over, I recorded four different ways the study could fail. You'll find them in the register. The one that matters most: satisfaction. I was not going to accept a fix that deprived me of coffee, which was recorded as assumption A8 in the register. There is no medical claim on trial here. This study counts cups and was never meant to replace the advice of my doctor.

Approach

I brewed every 16:1 cup by hand: 320g of water to 20g of medium-coarse ground coffee beans. Water pour was measured by weight and a timer for precision. That friction is the treatment; desirable difficulty as defined under Study parameters. I created a hands-free, voice-guided iOS Shortcut with timers and dictation checkpoints to train myself on the pour-over method. I also captured a 1–5 satisfaction score and a timestamp the moment I drank it with another Log a Cup hands-free iOS Shortcut. I spent the first two weeks of the study learning that I drank coffee mindlessly, then four weeks practicing this protocol, and watched for another two weeks to see if it stuck. I tracked BP with a physician escalation trigger set in advance at systolic over 150 or diastolic over 95. At the Week 4 checkpoint she cleared me to continue, and I dropped the daily readings: too noisy day to day, and the anxiety was not helping.

Research

The diary ran the length of the study; coded observation and the cross-validation log ran during the intervention weeks, 3–6, with eight of eight planned sessions completed. The phase summary reports full compliance: 56 of 56 diary days, zero cross-validation failures, 2.1% of optional fields missing. Beneath it, the log is less tidy: two checks failed and were resolved on the record, and the diary protocol notes one Week-4 recording reconstructed from memory. All 19 files are listed and downloadable, together or one at a time; the events file is a labelled reconstruction that reproduces the aggregates exactly, recorded under Open gaps.

Weekly record
Baseline (A1)Intervention (B)Washout (A2)Cups per week · diary-logged341617W1W2W3W4W5W6W7W8Satisfaction · 1–5, per cupguardrail 4.0registered floor 3.03.62
WeekPhaseCupsSatisfaction
W1Baseline (A1)303.73
W2Baseline (A1)343.65
W3Intervention (B)213.62
W4Intervention (B)184.33
W5Intervention (B)174.71
W6Intervention (B)164.81
W7Washout (A2)194.53
W8Washout (A2)174.47

Null hypothesis

Both reference lines come from the registered hypothesis: the ritual will cut consumption while maintaining satisfaction. The null is deprivation: drink less, enjoy it less. The 3.0 floor is the scale's neutral midpoint. If any week fell below it, the remaining cups would be worse than before the study, the null would stand, and PRE-001 fails. The 4.0 guardrail is the stricter test: baseline averaged 3.69, so holding 4.0 means the cups that remain beat baseline on average. Landing between the two lines would mean not falsified, but not a success either.

Falsification register

Scored in phase_summary.json against the registered thresholds. The events file is a labelled reconstruction that reproduces the aggregates exactly — see Open gaps.

IDFails ifThresholdActualResult
FC-001reduction <20%0.20.44PASS
FC-002satisfaction <3.033.62 lowestPASS
FC-003abandonment before Week 6complete Week 6completed all weeksPASS
FC-004washout rebound to baselineno return to baseline18 vs 32 baselinePASS

Baseline (A1)Weeks 1–2

Observation only

“The instrument runs; the treatment stays off” 64 cups in 14 days: 4.57 a day at satisfaction 3.69, with 66% of cups at work.

Intervention (B)Weeks 3–6

Active

“Every cup a pour-over decision” 72 cups in 28 days: 2.57 a day at satisfaction 4.37. Mastery rose from 0.14 to 1.0, and home cups from 27% to 57%.

Washout (A2)Weeks 7–8

Observation onlyHeld

“Does anything hold without the protocol?” 36 cups in 14 days: 2.57 a day at satisfaction 4.50. The pour-over ratio held at 0.42.

Assumption register
IDAssumptionCertaintyImportanceStatus
A1Coffee directly affects Bob's blood pressureHighCriticalExploratoryAssumed from medical guidance
A2The medication works best with reduced stimulant intakeMediumHighExploratoryAssumed from medical guidance
A3Bob drinks coffee out of habit not physiological needMediumHighValidatedBaseline tracking confirmed habitual pattern
A4Bob doesn't accurately track how much he drinksHighMediumValidatedSelf-report was 40% lower than actual
A5Coffee is available constantly (home work out)HighMediumObservedConfirmed during baseline
A6Coffee is a comfort ritual not just caffeine deliveryMediumHighValidatedObservation showed sensory engagement (smell > taste)
A7Bob wants to comply but lacks a satisfying alternativeLowCriticalValidatedPour-over became preferred method
A8Bob won't accept deprivation-based solutionsHighCriticalValidatedSatisfaction maintained above guardrail

Every assumption, scored for certainty and importance. Load-bearing premises that were never tested are flagged in their column.

Open gaps

  • diary_events.csv is a reconstructed export, labeled as such in the packet readme: the original device log was not retained, so the file was regenerated to reproduce every weekly aggregate exactly, on nominal 7-day weeks. The study’s weeks drifted; it ended on paper before Week 8 by that count. Both stand on the record.
  • PRE-001 registers a 3.0 satisfaction floor; a 4.0 guardrail appears packet-wide. Both stand. The stricter 4.0 governs the packet scoring.
  • DO-007 specifies daily continuous BP export; the decision log cut it at Week 4 (2024-01-29, “BP too noisy day-to-day”), so its status now reads Deprecated. Weekly monitoring under DO-005 continued: bp_log.csv carries a reading for every one of the eight weeks.
Mid-brew at the pour-over station: a stainless mesh filter loaded with grounds sits in the red ceramic dripper on a patterned cup, the Hario scale reading live beneath it, gooseneck kettle heated at the left
Intentional friction through ritual

The design is A–B–A: two weeks of baseline, four of intervention, two of washout, one subject throughout. A 14-day diary measured the baseline before anything changed. Every cup carried a satisfaction score, logged at the moment of drinking on a 1-to-5 scale. Eight observation sessions were recorded and coded, and the diary was cross-checked against them. Every threshold was registered before the first cup was logged (PRE-001).

measured All figures below are quoted from weekly_aggregates.csv and phase_summary.json. Not everything was contemporaneous: one Week-4 recording was reconstructed from memory the next day (04_diary_protocol.md), and one cup was added retroactively after a cross-validation catch (cross_validation_log.csv, 2024-01-29).

n = 1, and the participant is the experimenter: I knew every threshold while generating the data. No control, no blinding, no counterfactual; the comparison is me against myself in a different month. Observation is itself a treatment. Blood-pressure readings are exploratory monitoring only. Eight weeks is the full horizon.

01 Instrument

Brew a Cup

I designed a voice-paced Shortcut that walks the nine steps it takes to make a pour-over from beans and announces the running cup count.

Implemented

Design input
Reduce weekly cups
Actions
DO-001 Pour-over ritual protocolDO-006 Physician consultation triggers
Verification
Diary count comparison (baseline vs intervention)
Outcome
H-001PassAchieved 44% reduction (32→18)

Log a Cup

A spoken 1–5 satisfaction score captured at the moment of consumption, with context and method dictated hands-free.

Implemented

Design input
Maintain satisfaction
Actions
DO-002 Sensory engagement designDO-007 BP tracking instrumentation
Verification
Cup SEQ average calculation
Outcome
H-002PassFinal average 4.2; lowest point 3.4 (Week 3)

PRE-001: Fixed at Registration

Metric
cups_per_week
Succeeds at
≥20% reduction by Week 4

Fails if any of these hold

  1. FC-001reduction <20%
  2. FC-002satisfaction <3.0
  3. FC-003abandonment before Week 6
  4. FC-004washout rebound to baseline

Hypotheses registered before the study began

Registered
2024-01-05 · 09:00Z
First datum
2024-01-06 · 06:45Z
Outcome logged
2024-02-21 · 18:00Zcompleted

Open the evidence packet →

01 / 05 Registration · PRE-001

I fixed the hypothesis, the metric, the threshold and the four ways it could fail on 5 January. The first cup was logged the next morning. That order is the argument: choose a threshold after seeing the data and any result clears it, so the study can never fail. Choose it first and it can. The outcome was logged against those same criteria on 21 February, and the packet carries the timestamps.

PRE-001 · registered registered_at and outcome_logged_at quoted from pre_registrations.csv; the first diary event timestamp rendered live from diary_events.csv.

Baseline · cups per day

What I guessed
3–4
What the diary measured
4.57

The diary was the first instrument to fire, and it disagreed with me before anything else did.

Weeks 3–6 · diary against observer

  1. 01/15
  2. 01/17
  3. 01/22
  4. 01/24
  5. 01/29
  6. 01/31
  7. 02/05
  8. 02/07

24 of 26 reconciled. 01/29 cup count — diary 3, observed 1. Reviewed diary - found missed entry. 01/31 time accuracy — diary 280, observed 348. Recalibrated self-timing method.

02 / 05 Measurement · two loggers

Self-report is the weakest instrument in the study, so it never worked alone. I already knew it could be wrong: the baseline diary put my actual intake 40% above what I had guessed. During the intervention weeks, coded observation sessions logged the same behaviour independently and a cross-validation log reconciled the two. It caught a forgotten cup, added retroactively, and a self-timing method that needed recalibration. Both are on the record as failed checks.

A4 · measured Session and check counts from 05_observation_protocol.md and cross_validation_log.csv; the two failed checks are the log's passed=false rows (2024-01-29, 2024-01-31); the 40% figure is A4's validation note in 02_assumption_map.csv.
Baseline, weeks 1 to 2 32/wk
With the protocol, weeks 3 to 6 18/wk
Every support withdrawn, weeks 7 to 8 18/wk

03 / 05 Design · A–B–A

The washout answers the objection the design anticipates: anything works while you are paying attention to it. At Week 7 every active support was withdrawn: no protocol, no prompts. Consumption held: 19 cups in Week 7, 17 in Week 8, against a 32-cup baseline average. FC-004 asked whether the habit would rebound to baseline. It did not: 18.0 against 32.0.

DO-004 · FC-004 · measured Phase totals and the rebound comparison quoted from the phase summary; weekly washout counts from the weekly aggregates.

92.3%

of the drop came from the afternoon cups: the ones drunk out of habit rather than want. Friction removed them without being resisted.

Afternoon cups a week 13 1

Robert Duebelbeis, Applied ML Engineer & Product Designer
Didnt even try to cut these - too much work when im not really craving it.
  • 0.33 0.63 Supports it

    By Week 6 nearly two cups in three were made the hard way, and once nobody was asking, four in ten still were. Adoption outlived the instruction.

  • 3.62 4.81 Supports it

    It felt like punishment in Week 3 and like preference by Week 6, the dip-then-recovery desirable difficulty predicts.

  • 426s 248s Complicates it

    The ritual got 41.8% easier as mastery rose, so the friction largely went away, and consumption still did not climb back. Friction cannot be the whole mechanism.

04 / 05 Mechanism

Friction was supposed to be the mechanism, and for the first few weeks it was. Then it wore off. The ritual got easy, and the cups I had stopped drinking stayed gone. Something other than difficulty was holding the habit in place: preference, the ritual itself, or the plain fact that those afternoon cups were never really wanted. I did not design a way to tell which, so the study cannot say.

measured Mechanism indicators computed in phase_summary.json from weekly_aggregates.csv weeks 3–6; the quote is verbatim from 06_decision_log.csv (2024-02-07).

3.62

The lowest satisfaction the study recorded, against a floor of 3.0 fixed before any data existed. It happened in Week 3, the dip desirable difficulty predicts, and it is the closest this study came to failing.

Closest call

  • 44% vs 20% Cleared

    The registered bar was a fifth off the baseline. The run cleared it by more than double.

  • 18 vs 32 Cleared

    Withdrawal would have failed the study if consumption returned to the baseline. It stayed where the protocol had left it.

  • 8 of 8 weeks Cleared

    Binary: the study either ran to term or it did not. There is no margin to report, and pretending otherwise would invent one.

05 / 05 What the run delivered

This run delivered the record itself: four registered ways to fail, none of them hit. On this design, that result invites suspicion: n = 1, the participant is the experimenter, and I knew every threshold while generating the data. A clean sweep means the intervention worked, the hypotheses were too safe, or the measurement is broken; the last two cannot be ruled out from inside the study. Any method that never changes a decision is only being performed. The log shows ten decisions, none overturned, and no supersedes column to record one.

PRE-001 · recorded Outcome status and note quoted from pre_registrations.csv; decision count from 06_decision_log.csv: ten rows, headers without a supersedes column, no decision overturned.

Outcome

I closed the study on 2024-02-21 and logged the outcome against PRE-001, registered seven weeks earlier: a 20% reduction by Week 4 would count as success. The measured drop was 43.75%, from 32 cups a week to 18, and it held at 18 through both washout weeks with the protocol switched off. Satisfaction rose rather than fell: 3.69 at baseline, 4.37 under the protocol, 4.50 in washout, or 4.2 averaged across all 56 days. My doctor had wanted one cup a day; at the Week 6 call she accepted the sustained 44% instead. By Week 8 the decision log reads "18 cups/week holding. Actually prefer it now?" The question mark is mine. What surprised me was the instrument. Shortcuts run on any iPhone without an app build, so the diary was portable across an iOS population from day one, and voice pacing kept logging hands-free: usable mid-pour, wet-handed, without looking at a screen. That is the only reason I filed 56 of 56 diary days.

Measurement note

n=1, and the participant and the experimenter are the same person. There was no blinding and no control subject, so nothing here separates the protocol's effect from the effect of being watched by yourself. Consumption and satisfaction are self-reported through the diary instrument; the 56 of 56 filing rate is what makes them worth reading, not what makes them objective.

The 44% reduction is scored against PRE-001, registered on 5 January 2024 with a 20% threshold, seven weeks before the outcome was logged. Blood pressure was tracked only to a pre-set physician escalation trigger and was dropped after the Week 4 clearance; no cardiovascular claim is made or supported. This study counts cups.

Traceability

Six design inputs. Each names what satisfies it, how that was verified, and what the record says happened afterwards. The record shows five pass, one inconclusive.

6 of 6 items
IDHypothesisDesign OutputAcceptance CriteriaVerification MethodOutcome
H-001 Reduce weekly cups
≥30% reduction vs baselineDiary count comparison (baseline vs intervention) Pass
H-002 Maintain satisfaction
≥4.0 SEQ averageCup SEQ average calculation Pass
H-003 Habit transfers to washout
Consumption < baseline in weeks 7-8Washout consumption vs baseline comparison Pass
H-004 No adverse health events
Zero adverse events flaggedWeekly BP log + physician review Pass
H-005 BP correlation (exploratory)
Trend observation onlyTrend analysis of weekly BP readings Inconclusive
H-006 A hands-free ritual supports compliance
100% of intervention cups use guided shortcutBrew-log.csv row count vs diary pour-over count Pass
enesru