Skip to content

Interactive demonstration

The shutter

Two retention fixes enter the same A/B machinery. The fast fix cuts exit hazard by 15%, starting the day it ships. The slow fix is planted twice as strong — a 30% hazard cut — but its effect arrives gradually, over months, the way pacing and generosity changes do. Same game, same players, same statistics. The only thing you control is when the Thursday meeting reads the result. Companion to “The Experiment Ends Before the Player Does”

[ Illustrative synthetic data — demonstration only ]
fast fix — measured retention lift, with 95% CI band slow fix — measured retention lift, with 95% CI band zero — “no detectable effect”
power — probability the window detects the fast fix power — probability the window detects the slow fix
Thursday meeting, fast fix (−15% hazard, immediate)
Thursday meeting, slow fix (−30% hazard, arriving over τ)
first week the fast fix reads significant
first week the slow fix reads significant
week the measured ranking flips to the slow fix — before this, any n picks the wrong winner
the planted truth — which fix is actually stronger

What the window does

An experiment window is a filter with a shutter speed. Exit here is a weekly hazard: 5% of remaining players leave each week in control. The fast fix multiplies that hazard by 0.85 from day one — friction repairs behave this way, and by week four or five its retention lift stands clear of the confidence band. The slow fix multiplies hazard by 0.70 eventually, approaching that strength on the timescale τ — pacing, generosity, room-to-breathe changes behave this way, because they work through how players' relationship to the game drifts, and drift takes months. At the standard eight-week readout, the demo’s defaults hand down two verdicts: the weaker intervention ships, and the stronger one is shelved as “no detectable effect.” Slide the readout week forward and watch the second verdict change — same data-generating world, same planted truth, different Thursday.

The lower chart is the same fact as an instrument rating: statistical power — the probability the window detects each effect — as a function of window length. The fast fix’s power rises early. The slow fix’s power rises late, and both eventually fall: run the window long enough and attrition thins both arms until nothing is detectable — there is a horizon matched to each process, and the fiscal quarter is matched to neither. Sample size moves the power curves up — with enough players even the slow fix’s small early effect reads significant. What sample size never repairs is the ranking: at short windows the slow fix’s measured lift sits below the fast fix’s, because most of its effect hasn’t arrived yet, and the two curves swap order only at the crossover week the readout below tracks (about one and a half τ, on these defaults week 25). Until then, if the two fixes compete for one roadmap slot, the window picks the wrong winner at any n.

One honest note. The planted world is a cartoon — constant baseline hazard, one effect per arm, no seasonality, no interference between players. Every one of those simplifications favors the experimenter: real telemetry is noisier, and the slow effect is harder to see there, never easier. The direction of the bias this page demonstrates — evidence bases that systematically over-represent fast causes — survives every complication we removed. What a real diagnostic adds is the design work: holdouts and cohort structures that give slow processes an instrument matched to their timescale, stated before the data arrives.

Illustrative synthetic data — demonstration only. Model: weekly exit hazard h₀ = 0.05; fast arm h = h₀(1−0.15); slow arm h(t) = h₀(1−0.30·(1−e^{−t/τ})); retention S(W) = Π(1−h(t)); measured lift = difference in retention proportions with two-proportion 95% CI at n per arm; power = Φ(|Δ|/SE − 1.96). All curves computed exactly from these formulas — nothing is fitted.