A/B Test Significance Calculator

Enter visitors and conversions for each variant. Uses a two-proportion z-test to tell you whether the difference is real.

Why most A/B tests are called too early

With 40 clicks, a variant can look 50% better purely by chance. The z-test compares the difference between two conversion rates against the noise you would expect from random variation at that sample size. If the difference is small relative to the noise, the result is not significant and you should keep testing.

What the p-value means

The p-value is the probability of seeing a difference this large if the two variants were actually identical. A p-value of 0.03 means a 3% chance the gap is noise. At a 95% confidence level you need p < 0.05 to call a winner — and that still means roughly one in twenty "winners" is a false positive.

How much traffic you need

Required sample size scales with the inverse square of the effect you are trying to detect. Detecting a 10% relative lift needs roughly four times the traffic of detecting a 20% lift. As a rough guide for a 3% baseline conversion rate at 95% confidence:

Rules that keep tests honest