Track · 2:31 · Liner note
Variant B Was Kinder
An A/B test randomly splits users between two versions and compares a metric chosen in advance to see whether the difference is larger than chance would explain.
Split the users at random, show half of them version A and half version B, and compare a metric you chose in advance. That is an A/B test, and Variant B Was Kinder is the pleasant result: the difference is bigger than chance explains, and the new version treated users better.
Discipline keeps it honest. Decide the metric and sample size before you start, do not stop the moment the p-value dips below a threshold and check guardrail metrics for harm elsewhere. A result from a peeked test is a story, not a measurement.