A/B Test Significance Calculator

Two-proportion z-test: is the conversion difference between variant A and B statistically significant?

Loza CRM·Updated: October 2, 2026

Quick answer

Enter visitors and conversions for both variants — the calculator shows the p-value and whether the difference is statistically significant. p < 0.05 at 95% means the gap is real, not noise.

Variant A
Variant B
Variant B wins
CR A
10%
CR B
14%
Uplift B vs A
40%
z-score
2.7524
p-value
0.0059

How to use

  1. Enter visitors and conversions for variant A (control)
  2. Enter visitors and conversions for variant B (challenger)
  3. Pick a confidence level — 95% is the standard for media buying
  4. Read the verdict: significant or not, and which variant wins

Why you need it

An A/B test without a significance check is a coin flip with extra steps. A 2% conversion gap on 200 clicks is almost certainly noise; the same gap on 20,000 clicks is real. This calculator runs a two-proportion z-test and returns the p-value — the probability of seeing a difference this large if the variants were actually identical.

The most common mistake in affiliate testing is stopping early. Variant B pulls ahead after 300 visitors, you kill A — and the next thousand clicks flatten the difference completely. Peeking is not illegal, but acting on a lead before significance is how you burn the exact budget the test was meant to save.

Typical uses in media buying: two landing pages on a split test, two angles of the same offer, two prelanders before scaling. The discipline is simple — decide the confidence level before the test (95% is standard), run until you reach it, then make the call. If the result stays insignificant after a proper sample, the variants are equivalent for practical purposes: pick the cheaper one to run.

Worked example: landing A converts 100 of 1,000 visitors (10%), landing B converts 140 of 1,000 (14%). The uplift looks huge — 40% relative — and the calculator confirms it: p ≈ 0.006, significant at 95% and even 99%. But shrink the sample to 100 vs 100 visitors with 10 vs 14 conversions, and the same 40% uplift collapses into insignificance. That is the whole lesson: the effect size and the sample size are inseparable.

For campaigns running in Loza CRM, postback-attributed conversions feed the same math — this tool is for the quick check before you commit a week of spend to a winner that might be noise.

Loza CRM computes these metrics for your campaigns automatically — try it free.

Open Loza CRM

FAQ

What does p-value mean in plain terms?
The probability of seeing a difference this large by pure chance. p = 0.03 means: if the variants were identical, a gap this big would appear in 3% of tests — so at 95% confidence it counts as real.
How much traffic do I need for a valid test?
It depends on baseline conversion and effect size. At 2% CR, detecting a 20% relative difference takes thousands of visits per variant; detecting a 2× difference takes hundreds.
Does "not significant" mean the variants are equal?
No — it only means the data cannot distinguish the difference from noise yet. Usually it reads as "keep running", not "they are the same".
Which confidence level should I use?
95% is the standard. Use 99% for expensive decisions like replacing a main lander; 90% is acceptable for quick tests, but one in ten "wins" will be false.
Can I test more than two variants?
This tool compares pairs. For A/B/C run pairwise comparisons — and consider a stricter confidence level, since every extra comparison multiplies the chance of a false win.

Other tools