A/B Test Significance Calculator
Enter visitors and conversions for each variant. Uses a two-proportion z-test to tell you whether the difference is real.
Why most A/B tests are called too early
With 40 clicks, a variant can look 50% better purely by chance. The z-test compares the difference between two conversion rates against the noise you would expect from random variation at that sample size. If the difference is small relative to the noise, the result is not significant and you should keep testing.
What the p-value means
The p-value is the probability of seeing a difference this large if the two variants were actually identical. A p-value of 0.03 means a 3% chance the gap is noise. At a 95% confidence level you need p < 0.05 to call a winner — and that still means roughly one in twenty "winners" is a false positive.
How much traffic you need
Required sample size scales with the inverse square of the effect you are trying to detect. Detecting a 10% relative lift needs roughly four times the traffic of detecting a 20% lift. As a rough guide for a 3% baseline conversion rate at 95% confidence:
- Detect a 20% relative lift (3.0% → 3.6%): about 12,000 visitors per variant.
- Detect a 10% relative lift (3.0% → 3.3%): about 47,000 per variant.
- Detect a 5% relative lift (3.0% → 3.15%): about 185,000 per variant.
Rules that keep tests honest
- Fix the sample size in advance and do not peek daily to stop the moment it looks good — peeking inflates false positives dramatically.
- Run for whole weeks; weekday and weekend buyers behave differently.
- Test one change at a time, or you will not know what caused the effect.
- Use one-tailed tests only when you genuinely only care about improvement in one direction.