Free A/B testing tools: significance and sample size
Everything around an experiment: decide which hypothesis is worth testing, calculate the sample size and duration before launch, then read the result correctly with frequentist or Bayesian statistics. Formulas are shown, and results can be shared by link.
Tools in this collection
- RICE & ICE prioritization calculator
Which backlog items should go first by RICE or ICE score?
- Activation revenue calculator
How much revenue does a lift in activation rate bring?
- Funnel conversion calculator
Where do users drop off between steps, and which step is worth fixing first?
- A/B test calculator
Is the difference between A and B significant, and how many users does the test need?
- Survey sample size calculator
How many responses does a survey need for a given margin of error?
- Sample ratio mismatch calculator
Did the experiment really split traffic the way it was configured to?
When you need them
Before launching an A/B test
Rank hypotheses with ICE or RICE, find the funnel step with the largest loss, and estimate what the expected lift is worth in revenue.
Running and reading the test
Calculate the sample size and duration, then check significance, the confidence interval and probability to beat control. For surveys, size the sample by margin of error.
FAQ
How do I calculate A/B test significance?
Enter visitors and conversions for A and B. The calculator runs a two-proportion z-test and reports the p-value, relative lift with a confidence interval, and whether the difference is significant at your chosen level.
How many users does an A/B test need?
It depends on the baseline conversion rate, the minimum detectable effect, significance level and power. With a 5% baseline and a 10% relative lift at 95% significance and 80% power you need roughly 31,000 users per variant.
Frequentist or Bayesian — which should I use?
Frequentist tests control the false positive rate when you fix the sample size in advance. Bayesian mode gives the probability that B beats A and the expected loss, which is easier to explain. The calculator supports both.
Can I stop a test as soon as it becomes significant?
No. Peeking and stopping early inflates false positives. Run the test to the planned sample size, or use a method designed for sequential looks.
How do I choose which experiment to run first?
Score ideas with ICE or RICE, focus on the funnel step with the largest absolute loss, and estimate the revenue impact of the expected lift so the test is worth its traffic.
Is the statistics verified?
Yes. The distributions are checked against high-precision references, sample sizes against Evan Miller’s calculator, and the Bayesian probability against the closed-form formula and a Monte Carlo simulation.
Related collections
Product analytics
Funnel, activation, retention and revenue metrics with formulas and benchmarks.
Tools: 9Open →Growth
Unit economics, virality, acquisition links and conversion from trial to paid.
Tools: 13Open →UX research
Survey sample size and scoring for NPS, CSAT, CES, SUS and Kano surveys.
Tools: 8Open →