Business & Operations Tools

A/B Test Significance Calculator

Test whether a difference in conversion rates or means is statistically significant, with the confidence interval, effect size, achieved power and a warning before calling a winner.

  • p-value and CI
  • Relative lift
  • Power and peeking warnings
Runs in your browser

Everything you paste, type or drop is processed in this browser tab. It is not uploaded, logged, stored or sent to analytics.

Significance workspace

Examples:
What are you comparing?

1 Results per group

2 Test settings

Variants against control, times metrics you are judging it on. Alpha is split with Bonferroni.

3 What the data say

What the A/B Test Significance Calculator does

This calculator tests whether the difference between two versions in an A/B test is larger than chance would explain. For conversion rates it runs a two-proportion z-test; for averages such as order value it runs Welch's t-test. It reports the p-value, a confidence interval for the difference, an approximate interval for the relative lift, power for the effect you planned, and a check for sample ratio mismatch.

It deliberately never names a winner. It describes the strength of the evidence and the plausible size of the effect, and warns about the two habits that invalidate most tests: peeking and uncorrected multiple comparisons.

How to use it

  1. Choose conversion rates or averages.
  2. For rates, enter visitors and conversions for control and variant, plus the planned traffic split so the tool can check the split actually happened. For averages, enter the mean, standard deviation and number of observations in each group.
  3. Set alpha, the hypothesis (two-sided unless you decided otherwise before the test) and the number of comparisons - variants times metrics you are judging on.
  4. Enter the effect you planned for, to see whether the test had enough power to detect it.
  5. Read the verdict, the confidence interval and the working, then copy the summary into your test log.

Reading the results

A small p-value means data this extreme would be unusual if there were no real difference. It is not the probability that the variant is better, and it says nothing about whether the difference is large enough to matter.

The confidence interval is the more useful number: it is the range of effects consistent with the data. An interval from +0.15 pp to +1.85 pp says the effect is probably positive but could be tiny.

A non-significant result with low planned power is inconclusive, not negative. The test could not have reliably detected the effect you cared about.

Worked example: 200 of 1,000 against 250 of 1,000

Control converts 200 of 1,000 visitors (20%), the variant 250 of 1,000 (25%). The pooled rate is 450 / 2,000 = 22.5%, so the pooled standard error is sqrt(0.225 x 0.775 x (1/1,000 + 1/1,000)) = 0.018675 and z = 0.05 / 0.018675 = 2.677.

The two-sided p-value is 0.0074, and z squared (7.1685) equals the Pearson chi-square statistic without continuity correction, which is how you can cross-check it in any statistics package. The 95% interval for the difference, using the unpooled standard error of 0.018641, is +1.35 pp to +8.65 pp; the relative lift of 25% has an approximate interval of about 6% to 47%.

Formulas and scoring rules

Two-proportion z-test
z = (pB - pA) / sqrt(p(1-p)(1/nA + 1/nB)), p = (xA + xB) / (nA + nB)
p-value
two-sided 2(1 - Phi(|z|)); one-sided 1 - Phi(z) or Phi(z)Phi is the standard normal CDF (Hart / West algorithm, about 1e-14 accuracy).
CI of the difference
(pB - pA) +/- z(1 - alpha'/2) sqrt(pA(1-pA)/nA + pB(1-pB)/nB)Wald interval with unpooled standard error.
Relative lift CI (approximate)
exp(ln(pB/pA) +/- z sqrt((1-pA)/xA + (1-pB)/xB)) - 1Delta method on the log ratio; asymmetric, undefined with zero conversions.
Welch t-test
t = (mB - mA) / sqrt(sA^2/nA + sB^2/nB); df = (vA + vB)^2 / (vA^2/(nA-1) + vB^2/(nB-1))t distribution via the regularised incomplete beta function.
Sample ratio mismatch
chi2 = sum (observed - expected)^2 / expected, 1 df; flagged when p < 0.01
Multiple comparisons
alpha' = alpha / m (Bonferroni)p-values shown to 4 decimals.

Limitations: what the result does not prove

  • The tests assume each visitor is counted once and independently. Counting sessions or page views instead of users overstates the evidence.
  • Stopping as soon as the p-value first drops below alpha inflates false positives far beyond alpha. The results are only valid at the sample size you fixed in advance.
  • Statistical significance says nothing about practical importance, novelty effects that fade, or effects on metrics you did not test.
  • With very small counts the normal approximation is rough; an exact test is more appropriate there.

Privacy: where your data goes

Everything you paste, type or drop is processed in this browser tab. It is not uploaded, logged, stored or sent to analytics. Session recording and tag-manager scripts are switched off on this page.

Standards and sources

Frequently asked questions

How do I know if my A/B test is statistically significant?

Compare the p-value with your alpha, corrected for the number of comparisons. Below it, the data are unlikely under no difference. Always read the confidence interval as well, to see how large the effect plausibly is.

What does a p-value of 0.03 actually mean?

If there were truly no difference, results at least this extreme would occur about 3% of the time. It is not a 97% chance that the variant is better, and it does not measure the size of the effect.

Why is peeking at A/B test results a problem?

Each look is another chance for random noise to cross the threshold. Checking daily and stopping at the first significant result can push the real false-positive rate several times above alpha.

What is sample ratio mismatch?

A split of visitors that differs from the one you planned by more than chance allows. It usually points to a bug in assignment, redirects or tracking, and makes the conversion comparison untrustworthy until it is explained.

When should I use a t-test instead of a z-test?

Use the proportion z-test for yes/no outcomes such as conversion. Use Welch's t-test for continuous values such as revenue per visitor or time on task, where you have means and standard deviations.

Why is the relative lift interval not symmetric?

A ratio cannot fall below minus 100% but can rise without limit, so its uncertainty is naturally skewed. The interval is computed on the log scale and transformed back, which reflects that.

My result is not significant - does that mean there is no difference?

No. It means the data do not rule out zero. If power for your planned effect was low, the test simply could not detect it; the interval shows which effects remain plausible.

Last reviewed by the A2Z.Tools team against the sources listed above.

Rate this tool

Was this tool useful? Your feedback helps us improve it.

No ratings yet — be the first to rate this tool.
Your rating (required)
0 / 2000

Please do not include passwords, payment details or other sensitive information.

Your feedback is sent privately to the A2Z.Tools team and will not be posted publicly.

Add this calculator to your website. Free, responsive, no ads, no sign-up - copy one line of code.

Embed this tool