A Great Place to Upskill
Get the latest updates from Product Space
How to know whether your change actually worked, and tell a real result from noise
Learn to design, run, and honestly read A/B tests: hypotheses, sample size, guardrail metrics, and the six ways tests quietly lie.
The three questions that kill half your test ideas
When to just ship it
When to test, and when to go and talk to someone instead
The cost of a test, honestly counted
Half of all test ideas shouldn't be tested. This chapter is about finding out which half, before you build anything.
The four ways a test fails before it starts
What a test can and cannot tell you
Why "we'll just test it" is often the wrong answer
What this course covers and what it doesn't
Four ways a test is doomed before it starts, what a test can actually tell you, and why "let's just test it" is often the wrong instinct.
Why "we think this will improve conversion" isn't a hypothesis
The four-part hypothesis
Saying in advance what would make you wrong
One change or several: the honest trade-off
Why a hypothesis needs to be specific enough to fail, and how to write one that produces a real decision instead of a debate.
The primary metric, and why there's exactly one
Guardrail metrics: what you must not break to win
Metrics that can be won by making the product worse
Leading and lagging: what you can actually measure in three weeks
One primary metric, a few guardrails, and an honest look at whether your metric can be won by making the product worse.
Sample size without the statistics degree
The four numbers that decide everything
Why most tests are too small to detect what you're hoping for
The weekly cycle rule, and why you never stop on a Tuesday
The arithmetic that decides whether your test is capable of answering your question, explained without a statistics degree.
The setup checklist before you turn it on
The A/A test, and why it's worth the week
What to check on day two
The six things that invalidate a test mid-flight
The setup checks that prevent a wasted three weeks, and the six things that invalidate a test mid-flight.
What "significant" actually means, in plain words
The flat result: the most common and most misread outcome
Effect size: the number that matters more than the p-value
When the result contradicts what you believed
What the numbers actually mean, how to handle the flat result, and what to do when the answer isn't the one you wanted.
Winning isn't the same as worth shipping
The decision the test cannot make for you
Rolling out, and watching for the effect to fade
Writing it up so the next person benefits
Winning isn't the same as worth shipping. How to decide, roll out safely, and write up results so the next person benefits.
Peeking, and why it's the most common error
Slicing until something is significant
Novelty and primacy effects
Contamination, sample mismatch, and survivorship
Each of these produces a confident, wrong answer, and each one is committed by good teams with good intentions.
Five situations where testing won't help
Too little traffic: what to do instead
Changes too big to isolate
Knowing when to decide without a test
The most valuable judgement in this course: knowing when to test, and when to do something else entirely.
Part 1: A hypothesis, metrics and a sample size calculation
Part 2: A run, with a mid-flight check and no peeking
Part 3: A readout with a decision and an honest limits section
One properly designed, properly run, honestly reported A/B test, applying everything from the course to a real test on a real product.
Before you build: deciding whether to test, writing a hypothesis, choosing metrics, sample size and duration
Running it: pre-launch review
Reading it: reading a result honestly, auditing for the six lies
Afterwards: the test readout
Checking yourself: is testing right for your situation
Every copy-and-use prompt from the course, organized by the stage of testing it supports.
Plain-language definitions of every term used in the course, from A/A test to underpowered, for quick reference while working.