Clew

HomeToolsA/B testingA/B test calculator

A/B test calculator with significance, sample size and duration

Check whether a difference in conversion rates is significant, find out how many visitors and days a test needs, and look at the result through Bayesian statistics. Distribution charts with confidence intervals and percentiles, like the classic calculators.

● Free, no sign-upUpdated:

A — control

B — variant

Confidence level
Hypothesis
Two-sided catches both a lift and a drop. Choose one-sided in advance, before the test starts.

—

 

Conversion A— 
Conversion B— 
Difference— 
Relative lift— 
p-value— 
Power— 
Conversion rate distributions of A and Bshaded — confidence interval
Conversion rate distributions of A and B
A — controlB — variantdashed — interval bounds, solid — median
Distribution percentiles
2.5th percentile50th (median)97.5th percentile
A———
B———
B − A———
Distribution of the difference B − A, ppinterval clear of zero — significant
Distribution of the difference B − A, pp

Don’t peek. Decide once, when the sample size you calculated in advance is reached. Checking the p-value daily and stopping at the first “significant” multiplies the false positive rate.

Everything is calculated in your browser — nothing you enter is sent anywhere.

How to use it

  1. Pick a modeBefore launch, use “Sample size & duration” to learn how many visitors and days you need. After the test, use “Check a result” or “Bayesian”.
  2. Enter your dataVisitors and conversions for control (A) and variant (B). A conversion is a user who completed the goal, not an event count.
  3. Read the verdictThe verdict at the top explains the result in words; below are the p-value, intervals and distribution charts. Hover a chart to see the value and percentile.
  4. Share the linkInputs are saved in the page URL. “Copy link” gives a URL where a teammate sees exactly the same calculation.

How the calculator tests significance

“Check a result” uses the classic two-proportion z-test, the same test most A/B calculators use. The null hypothesis: conversion rates of A and B are equal, and the observed difference is sampling noise.

First the rates pA = xA / nA and pB = xB / nB are computed. For the test, the standard error of the difference uses the pooled rate — under the null hypothesis both groups share one conversion rate:

p̂ = (x_A + x_B) / (n_A + n_B) SE₀ = √( p̂ · (1 − p̂) · (1/n_A + 1/n_B) ) z = (p_B − p_A) / SE₀ p-value = 2 · (1 − Φ(|z|)) — two-sided test p-value = 1 − Φ(z) — one-sided, H₁: p_B > p_A

Φ is the standard normal CDF. The difference is significant when the p-value is below α = 1 − confidence level.

Confidence intervals

Intervals are built without pooling — each group keeps its own variance (the Wald interval):

SE_A = √( p_A (1 − p_A) / n_A ) CI_A = p_A ± z₁₋α/₂ · SE_A SE_Δ = √( SE_A² + SE_B² ) CI_Δ = (p_B − p_A) ± z₁₋α/₂ · SE_Δ

The interval for the relative lift pB/pA − 1 uses the delta method. These normal distributions N(p, SE) are what the charts show: the shaded area is the confidence interval, dashed lines are its bounds (the 2.5th and 97.5th percentiles at 95%), the solid line is the median.

Power and sample checks

“Power” is observed (post-hoc) power: the probability that a test of this size would detect a difference equal to the observed one. It is a hint about whether you have enough data, not evidence for the hypothesis — observed power is a direct function of the p-value and adds no new information.

The calculator also checks for SRM (sample ratio mismatch) with a χ² test for an even traffic split. If visitors are split noticeably unevenly, the randomization is most likely broken and comparing conversion rates is meaningless.

Sample size and test duration

How many visitors you need depends on four numbers: the baseline conversion, the minimum detectable effect (MDE), the significance level α and power 1 − β. The calculator uses the normal-approximation formula for comparing two proportions — the same one as Evan Miller’s calculator, so the results match:

δ = MDE (absolute; a relative MDE is multiplied by p) n = ( z₁₋α/₂ · √(2p(1 − p)) + z₁₋β · √(p(1 − p) + (p + δ)(1 − p − δ)) )² / δ²

n is the number of visitors in each variant. Test duration = n × number of variants ÷ visitors per day.

With more than two variants, every comparison with control is another chance to be wrong. The calculator applies the Bonferroni correction: α is divided by the number of comparisons. It is conservative but simple and honest; the price is a larger sample.

Rules of thumb:

  • Run tests in whole weeks — at least 7 days, ideally 14: weekday and weekend audiences differ.
  • Halving the MDE needs roughly four times the traffic. If a test would take six months, test a bolder change or a metric closer to the action.
  • For onboarding, metrics with a high baseline — step completion, the first key action — work better than payment: they need several times fewer visitors.

Bayesian approach: the probability that B is better

Bayesian mode answers the question people usually ask: how likely is B to be better than A? Each variant’s conversion rate is described by a beta distribution. With a Beta(α, β) prior and x conversions out of n visitors, the posterior is Beta(α + x, β + n − x) — the conjugate beta-binomial model.

P(B > A) = ∫ f_B(x) · F_A(x) dx Risk of choosing B = E[ max(p_A − p_B, 0) ] Risk of keeping A = E[ max(p_B − p_A, 0) ]

The probability to beat control is computed by numerical integration and checked in our tests against Evan Miller’s closed-form formula for integer parameters and against a Monte Carlo simulation. Expected loss is how much conversion you lose on average if the choice turns out wrong. A convenient stopping rule: end the test when the risk of the chosen variant is below a threshold that does not matter for the business.

Bayesian inference does not excuse sloppy process: peeking and stopping “at the first 95%” hurt it too. The uniform Beta(1, 1) prior barely affects the result once you have hundreds of conversions; an informative prior only makes sense if it honestly reflects past tests.

How to read the result and which mistakes to avoid

What “statistically significant” means

A p-value of 0.03 means: if there were really no difference, a gap this large or larger would appear in about 3% of tests. It is not the probability that B is better, and not the size of the effect. The confidence interval shows the size: “+0.2 to +1.4 pp” is more honest than “+16%, significant”.

Common mistakes

  • Peeking. Checking significance daily and stopping at the first p < 0.05 is a reliable way to get a false win. Calculate the sample size in advance and decide once.
  • Stopping at “significant” before a full week. Novelty effects and day-of-week patterns distort the first days of a test.
  • Too many metrics and segments. Check twenty metrics and you will find one “significant” by chance. Choose the primary metric before launch; treat segments as ideas for the next test.
  • Uneven split (SRM). If group sizes diverge more than chance allows, look for a bug in the traffic split, not for a winner.
  • Counting events as conversions. A user who clicks five times is one conversion. The test assumes independent observations.
  • “Not significant” = “no effect”. A non-significant result on a small sample only means there was not enough data. Look at the interval: if it spans both −5% and +10%, the test decided nothing.

If you are testing onboarding — a tour versus no tour, or two checklist versions — compare activation (the first key action), not “completed the tour”. How to raise it: How to increase user activation in SaaS.

Sources

  1. Fleiss J. L., Levin B., Paik M. C. Statistical Methods for Rates and Proportions. 3rd ed. Wiley, 2003 — two-proportion z-test and sample size.
  2. Miller E. Sample Size Calculator — evanmiller.org/ab-testing/sample-size.html; Miller E. Formulas for Bayesian A/B Testing — evanmiller.org/bayesian-ab-testing.html.
  3. Kohavi R., Tang D., Xu Y. Trustworthy Online Controlled Experiments. Cambridge University Press, 2020 — SRM, peeking, choosing metrics.
  4. Marsaglia G. Evaluating the Normal Distribution. Journal of Statistical Software, 11(4), 2004 — computing Φ(x).
  5. Press W. H. et al. Numerical Recipes. 3rd ed., §6.4 — the incomplete beta function.

FAQ

How do I check the statistical significance of an A/B test?

Enter visitors and conversions for both groups in “Check a result”. The calculator runs a two-proportion z-test and returns the p-value: if it is below 1 − confidence level (for example 0.05 at 95%), the difference is statistically significant. Also look at the confidence interval of the difference — it shows how large the effect is.

How long should an A/B test run?

Until it reaches the sample size calculated in advance, and no less than 7 full days. In “Sample size & duration”, enter the baseline conversion, the minimum effect and daily traffic — the calculator returns visitors per variant and the duration in days.

Which confidence level should I use: 90, 95 or 99%?

The standard is 95%. 90% suits cheap, reversible changes where a false win costs little; 99% suits risky decisions such as pricing. The higher the confidence level, the larger the sample you need.

What is the difference between a one-sided and a two-sided test?

A two-sided test checks whether B differs from A in either direction; a one-sided test only checks whether B is better. One-sided needs less data, but you must choose it before launch and it will not flag a drop. When in doubt, use two-sided.

What is MDE?

The minimum detectable effect is the smallest difference in conversion that the test reliably detects (with the chosen power). Set it before the test: the smaller the MDE, the more visitors you need — roughly in inverse proportion to the square of the effect.

Do the results match Evan Miller’s calculator?

Yes. Sample size uses the same formula; our automated tests compare the result with his calculator on several parameter sets, with a difference of at most one visitor due to the precision of the inverse normal function.

Is the Bayesian approach better than the frequentist one?

It is not better; it answers a different question. A frequentist test tells you how surprising the data are if there is no difference; a Bayesian one tells you how likely B is to be better and how much you risk losing. The second is easier to explain to the business, but both need stopping rules and discipline.

Can I test more than two variants?

Yes. Set the number of variants in “Sample size & duration”: the calculator applies the Bonferroni correction and increases the sample. When checking the result, compare each variant with control separately, using α divided by the number of comparisons.

More in A/B testing

Open the collection →

Tours, tooltips and checklists whose impact shows up in the numbers.

Try Clew for freeHow it worksFree plan forever · no credit card

Product tours, popups and onboarding in the age of AI

7 patterns with step counts, copy rules and what to measure, what AI changes, and a checklist before you publish. PDF, 2 pages. What is inside →