Variants

The first row is your control (baseline) — every other variant is compared against it. Add as many as you're testing.

# Name Visitors Conversions

Confidence level

Control conversion rate
Variant Conv. rate Uplift Z-score P-value Result

About this calculator

This uses a two-proportion z-test, the standard method for comparing conversion rates between two groups. Each variant is tested individually against the control: pooling the two groups' conversions to estimate a shared standard error, then measuring how many standard errors apart the two observed rates are (the z-score).

The p-value is the probability of seeing a difference this large (or larger) if there were actually no real difference between that variant and the control. If the p-value is below your chosen significance threshold (5% at 95% confidence, 1% at 99%, and so on), the result is considered statistically significant.

Testing several variants against the same control at once increases the chance that at least one shows up "significant" purely by chance — known as the multiple comparisons problem. Treat borderline results across many variants with extra caution, or raise your confidence level accordingly.

A significant result doesn't guarantee the uplift is real forever — it means the difference is unlikely to be due to random chance alone, given the sample sizes you tested. Small sample sizes need much larger differences to reach significance than large ones do.

Frequently asked questions

How many visitors do I need for an A/B test to be significant?

There is no single number — it depends on your baseline conversion rate and how small a difference you want to detect. Smaller differences need far bigger samples, so detecting a change of a percentage point or two on a typical conversion rate usually takes thousands of visitors per variant.

What does a p-value of 0.05 actually mean?

It means there is roughly a 5% chance of seeing a difference at least this large if the variants were genuinely identical. When the p-value falls below your threshold — 5% at 95% confidence, 1% at 99% — the result is called statistically significant.

Can I stop the test as soon as it shows significance?

Better not to. Checking repeatedly and stopping the moment a result turns significant sharply increases false positives. Decide roughly how long you will run the test before you start, and let it finish.

More Business & Analysis calculators
🤖
AI Cost
Token prices, caching and batch discounts → what an LLM API really costs.
📊
Margin vs Markup
Cost + margin% or markup% → sell price, and back-solve either way.
📈
YoY / MoM Growth
Feed in a monthly series, get month-on-month and year-on-year change.
🧾
Tax / VAT
Add or remove VAT, plus a progressive income tax band split.