Study Parameters
Sole author and sole subject of an n=1 self-experiment: I designed the protocol, built the logging instrument, and was the participant. My doctor put me on blood-pressure medication and told me to cut back on coffee. The reason she gave was work: the stress of the job I was doing at the time, and what it was doing to my blood pressure. Coffee was the lever she could name. She did not say by how much. Heck, I didn't even know how much I was really drinking. All I knew is that I guessed around 3–4 cups every day. But before changing anything, I captured a true baseline by writing down every cup I drank for two weeks. The answer was closer to 32 a week; 4.57/day, 40% above my own estimate. "Way more than I thought," the decision log says. I wanted to cut back, but I didn't want to replace coffee with something worse. So this study pulled a different lever: not a limit on cups I could have, but a harder way to make each one.
Before the first pour-over, I recorded four different ways the study could fail. You'll find them in the register. The one that matters most: satisfaction. I was not going to accept a fix that deprived me of coffee, which was recorded as assumption A8 in the register. There is no medical claim on trial here. This study counts cups and was never meant to replace the advice of my doctor.
I brewed every 16:1 cup by hand: 320g of water to 20g of medium-coarse ground coffee beans. Water pour was measured by weight and a timer for precision. That friction is the treatment; desirable difficulty as defined under Study parameters. I created a hands-free, voice-guided iOS Shortcut with timers and dictation checkpoints to train myself on the pour-over method. I also captured a 1–5 satisfaction score and a timestamp the moment I drank it with another Log a Cup hands-free iOS Shortcut. I spent the first two weeks of the study learning that I drank coffee mindlessly, then four weeks practicing this protocol, and watched for another two weeks to see if it stuck. I tracked BP with a physician escalation trigger set in advance at systolic over 150 or diastolic over 95. At the Week 4 checkpoint she cleared me to continue, and I dropped the daily readings: too noisy day to day, and the anxiety was not helping.
The diary ran the length of the study; coded observation and the cross-validation log ran during the intervention weeks, 3–6, with eight of eight planned sessions completed. The phase summary reports full compliance: 56 of 56 diary days, zero cross-validation failures, 2.1% of optional fields missing. Beneath it, the log is less tidy: two checks failed and were resolved on the record, and the diary protocol notes one Week-4 recording reconstructed from memory. All 19 files are listed and downloadable, together or one at a time; the events file is a labelled reconstruction that reproduces the aggregates exactly, recorded under Open gaps.
| Week | Phase | Cups | Satisfaction |
|---|---|---|---|
| W1 | Baseline (A1) | 30 | 3.73 |
| W2 | Baseline (A1) | 34 | 3.65 |
| W3 | Intervention (B) | 21 | 3.62 |
| W4 | Intervention (B) | 18 | 4.33 |
| W5 | Intervention (B) | 17 | 4.71 |
| W6 | Intervention (B) | 16 | 4.81 |
| W7 | Washout (A2) | 19 | 4.53 |
| W8 | Washout (A2) | 17 | 4.47 |
Null hypothesis
Both reference lines come from the registered hypothesis: the ritual will cut consumption while maintaining satisfaction. The null is deprivation: drink less, enjoy it less. The 3.0 floor is the scale's neutral midpoint. If any week fell below it, the remaining cups would be worse than before the study, the null would stand, and PRE-001 fails. The 4.0 guardrail is the stricter test: baseline averaged 3.69, so holding 4.0 means the cups that remain beat baseline on average. Landing between the two lines would mean not falsified, but not a success either.
Falsification register
Scored in phase_summary.json against the registered thresholds. The events file is a labelled reconstruction that reproduces the aggregates exactly — see Open gaps.
| ID | Fails if | Threshold | Actual | Result |
|---|---|---|---|---|
| FC-001 | reduction <20% | 0.2 | 0.44 | PASS |
| FC-002 | satisfaction <3.0 | 3 | 3.62 lowest | PASS |
| FC-003 | abandonment before Week 6 | complete Week 6 | completed all weeks | PASS |
| FC-004 | washout rebound to baseline | no return to baseline | 18 vs 32 baseline | PASS |
Baseline (A1)Weeks 1–2
Observation only
“The instrument runs; the treatment stays off” 64 cups in 14 days: 4.57 a day at satisfaction 3.69, with 66% of cups at work.
Intervention (B)Weeks 3–6
Active
“Every cup a pour-over decision” 72 cups in 28 days: 2.57 a day at satisfaction 4.37. Mastery rose from 0.14 to 1.0, and home cups from 27% to 57%.
Washout (A2)Weeks 7–8
Observation onlyHeld
“Does anything hold without the protocol?” 36 cups in 14 days: 2.57 a day at satisfaction 4.50. The pour-over ratio held at 0.42.
| ID | Assumption | Certainty | Importance | Status |
|---|---|---|---|---|
| A1 | Coffee directly affects Bob's blood pressure | High | Critical | ExploratoryAssumed from medical guidance |
| A2 | The medication works best with reduced stimulant intake | Medium | High | ExploratoryAssumed from medical guidance |
| A3 | Bob drinks coffee out of habit not physiological need | Medium | High | ValidatedBaseline tracking confirmed habitual pattern |
| A4 | Bob doesn't accurately track how much he drinks | High | Medium | ValidatedSelf-report was 40% lower than actual |
| A5 | Coffee is available constantly (home work out) | High | Medium | ObservedConfirmed during baseline |
| A6 | Coffee is a comfort ritual not just caffeine delivery | Medium | High | ValidatedObservation showed sensory engagement (smell > taste) |
| A7 | Bob wants to comply but lacks a satisfying alternative | Low | Critical | ValidatedPour-over became preferred method |
| A8 | Bob won't accept deprivation-based solutions | High | Critical | ValidatedSatisfaction maintained above guardrail |
Every assumption, scored for certainty and importance. Load-bearing premises that were never tested are flagged in their column.
Open gaps

The design is A–B–A: two weeks of baseline, four of intervention, two of washout, one subject throughout. A 14-day diary measured the baseline before anything changed. Every cup carried a satisfaction score, logged at the moment of drinking on a 1-to-5 scale. Eight observation sessions were recorded and coded, and the diary was cross-checked against them. Every threshold was registered before the first cup was logged (PRE-001).
n = 1, and the participant is the experimenter: I knew every threshold while generating the data. No control, no blinding, no counterfactual; the comparison is me against myself in a different month. Observation is itself a treatment. Blood-pressure readings are exploratory monitoring only. Eight weeks is the full horizon.
Brew a Cup
I designed a voice-paced Shortcut that walks the nine steps it takes to make a pour-over from beans and announces the running cup count.
Implemented
Log a Cup
A spoken 1–5 satisfaction score captured at the moment of consumption, with context and method dictated hands-free.
Implemented
I closed the study on 2024-02-21 and logged the outcome against PRE-001, registered seven weeks earlier: a 20% reduction by Week 4 would count as success. The measured drop was 43.75%, from 32 cups a week to 18, and it held at 18 through both washout weeks with the protocol switched off. Satisfaction rose rather than fell: 3.69 at baseline, 4.37 under the protocol, 4.50 in washout, or 4.2 averaged across all 56 days. My doctor had wanted one cup a day; at the Week 6 call she accepted the sustained 44% instead. By Week 8 the decision log reads "18 cups/week holding. Actually prefer it now?" The question mark is mine. What surprised me was the instrument. Shortcuts run on any iPhone without an app build, so the diary was portable across an iOS population from day one, and voice pacing kept logging hands-free: usable mid-pour, wet-handed, without looking at a screen. That is the only reason I filed 56 of 56 diary days.
n=1, and the participant and the experimenter are the same person. There was no blinding and no control subject, so nothing here separates the protocol's effect from the effect of being watched by yourself. Consumption and satisfaction are self-reported through the diary instrument; the 56 of 56 filing rate is what makes them worth reading, not what makes them objective.
The 44% reduction is scored against PRE-001, registered on 5 January 2024 with a 20% threshold, seven weeks before the outcome was logged. Blood pressure was tracked only to a pre-set physician escalation trigger and was dropped after the Week 4 clearance; no cardiovascular claim is made or supported. This study counts cups.
Six design inputs. Each names what satisfies it, how that was verified, and what the record says happened afterwards. The record shows five pass, one inconclusive.
| ID | Hypothesis | Design Output | Acceptance Criteria | Verification Method | Outcome |
|---|---|---|---|---|---|
| H-001 | Reduce weekly cups | ≥30% reduction vs baseline | Diary count comparison (baseline vs intervention) | Pass | |
| H-002 | Maintain satisfaction | ≥4.0 SEQ average | Cup SEQ average calculation | Pass | |
| H-003 | Habit transfers to washout | Consumption < baseline in weeks 7-8 | Washout consumption vs baseline comparison | Pass | |
| H-004 | No adverse health events | Zero adverse events flagged | Weekly BP log + physician review | Pass | |
| H-005 | BP correlation (exploratory) | Trend observation only | Trend analysis of weekly BP readings | Inconclusive | |
| H-006 | A hands-free ritual supports compliance | 100% of intervention cups use guided shortcut | Brew-log.csv row count vs diary pour-over count | Pass |