Clew

HomeToolsA/B testingSRM calculator

Sample ratio mismatch calculator for A/B tests

Before you read any result of an experiment, check that the users arrived in the proportions you configured. Enter the intended split and the observed users per variant — two variants or ten — and get the chi-square goodness-of-fit test, the p-value, and the number of users that would have to move to make the split exact.

● Free, no sign-upUpdated:

Variants

    The split can be anything proportional: 50 / 50, 1 / 1, 34 / 33 / 33 or 0.1 / 0.9 all work — the values are normalised. Users must be whole numbers, counted the same way for every variant and at the same point in the funnel.

    p-value threshold
    Experimentation platforms use 0.0005, not 0.05: an SRM check runs on every experiment, every day, so a 5% false-alarm rate would cry wolf constantly. The other two are here to show how much the verdict depends on the line you draw.

    —

     

    p-value— 
    Chi-square— 
    Largest deviation— 
    Misplaced users— 
    Detectable imbalance— 
    Users in total— 
    Configured split against observed split
    Configured split against observed split
    configuredobserved

    Per variant

    Configured share, observed share, expected users, difference and standardised residual for each variant
    VariantConfiguredObserved Expected usersDifferenceResidual

    Method. Pearson chi-square goodness of fit against the configured split, with k − 1 degrees of freedom for k variants, and the exact upper tail of the chi-square distribution (no normal approximation). The standardised residual of a variant is its difference divided by the square root of its expected count; squared, it is that variant’s share of the chi-square statistic, which is how you find the culprit when more than two variants are involved.

    Everything is calculated in your browser — nothing you enter is sent anywhere.

    How to use it

    1. Take the counts from one placeUsers per variant from the same query, the same date range and the same funnel step — ideally the exposure event your experiment platform logs, not a downstream page view.
    2. Enter the split you configuredWhat the experiment was set up to do, not what it did. If you ramped from 5/95 to 50/50 mid-test, the check is meaningless for the whole period: run it on each period separately.
    3. Read the p-value, not the percentagesA 0.4% gap is nothing in 2,000 users and an emergency in 2,000,000. The p-value is what accounts for that; the threshold is 0.0005 because you run this check constantly.
    4. If it fails, find the cause before anything elseThe residual column names the variant to start from. Work through assignment, redirects, bot filtering, caching and the analytics join — the causes checklist below is in the order that finds most problems fastest.

    How the SRM check works

    A sample ratio mismatch (SRM) is a difference between the traffic split an experiment was configured with and the split that actually shows up in the data, large enough that chance is not a plausible explanation. It is the single most useful trust check in online experimentation, because almost anything that can silently corrupt an experiment also disturbs the ratio.

    expected_i = N · w_i / Σw configured share of variant i χ² = Σ (observed_i − expected_i)² / expected_i df = k − 1 for k variants p = P(χ² > observed χ²) upper tail, k − 1 df residual_i = (observed_i − expected_i) / √expected_i

    With two variants at 50/50 this reduces to the familiar one-degree-of-freedom test. With more variants the chi-square says only that something is off; the residuals say which variant, because each residual squared is that variant’s contribution to the total.

    Why the threshold is 0.0005 and not 0.05

    An SRM check is not a one-off hypothesis test. It runs on every experiment, often every day, so at p < 0.05 roughly one healthy experiment in twenty would be flagged, every day, and the alarm would be ignored within a week. The convention that came out of running this at scale — Fabijan and colleagues at Microsoft, and the same practice at other large platforms — is p < 0.0005. That trades sensitivity for an alert people still believe, which is the right trade for a guardrail.

    The flip side matters too: a small experiment cannot detect a small mismatch at that threshold. The “detectable imbalance” figure above tells you how far off the split would have had to be for your traffic to notice, so a pass on 900 users is not the same statement as a pass on nine million.

    What causes SRM

    1. Assignment. Users bucketed more than once, a hash that is not uniform, a variant that errors during assignment, or a default that quietly sends failures to control.
    2. Exposure logging. The variant that renders slower loses the users who leave before the exposure event fires — a redirect test is the classic case, and it usually makes the treatment look smaller and better.
    3. Filtering applied after assignment. Bot rules, region rules, consent gates or “activated users only” filters that correlate with the variant.
    4. Caching and CDNs. A cached page served to the wrong bucket, or a variant excluded from the cache and therefore from a chunk of traffic.
    5. The analytics join. Deduplication, session stitching, late-arriving events, or joining exposure to a table with different retention.
    6. Ramping and targeting changes during the window, which are not bugs but do invalidate a check over the whole period.

    When this check misleads you

    SRM is a good guardrail and a bad oracle. These are the ways it is misread.

    • A pass does not mean the experiment is trustworthy. It means the totals are not visibly wrong. An experiment can have a perfect ratio and still be broken by an instrumentation bug that hits both variants, a metric defined differently per variant, or a bucketing scheme that splits evenly but not randomly.
    • It is not a test of the experiment’s result. The p-value here says nothing about whether your change worked. It is entirely separate from the significance of your metrics — and if it fails, the significance of your metrics is unreadable, however small their p-value is.
    • Checking every day inflates false alarms. The 0.0005 threshold exists because of this, and it is still not a licence to peek at will. If you check a hundred healthy experiments daily for a month, expect the occasional flag; look for a mismatch that persists and grows, not one that appears once.
    • An uneven split is not automatically a mismatch. If you configured 90/10, enter 90/10. Most “SRM calculators” assume an even split, which turns every ramped or holdback experiment into a false alarm.
    • Checking a filtered population usually creates the mismatch. Run the check on the exposed population. Counting only users who reached checkout, or only paying users, introduces exactly the variant-correlated filtering that SRM is supposed to detect.
    • Chi-square needs enough expected users per variant. Below roughly five expected per variant the p-value is arithmetic, not evidence. This page says so when it happens.
    • Overall pass, segment fail. A mismatch can live in one platform, one country or one browser and cancel out in the total. If a bug is suspected, check the ratio inside segments as well — and remember that testing many segments brings back the multiple-comparison problem.
    • Fixing the numbers is not fixing the cause. Reweighting, dropping the extra users or “rerunning the analysis on the matched sample” hides the symptom. The users who went missing are not a random sample of your traffic, which is the whole reason SRM invalidates a result.

    The right response to a mismatch is to stop reading metrics, find the mechanism, fix it, and rerun. That feels expensive, and it is much cheaper than shipping a change based on a broken experiment.

    After the check

    • If it passed, read the result properly. The A/B test calculator gives significance, confidence intervals and a Bayesian view, and warns about a two-variant ratio imbalance as you go.
    • If the test is still running, know what it can prove. The same calculator’s sample size and minimum detectable effect will tell you whether the experiment can settle the question at all.
    • If the experiment is about onboarding, check the step before you check the stats. The funnel conversion calculator finds which step leaks most, which is usually where a test is worth running.
    • Watch what the winner does to revenue. The activation rate revenue calculator turns a lift in activation into MRR, which is the number that decides whether the change was worth the release.

    Sources

    1. Fabijan A., Gupchup J., Gupta S., Omhover J., Qin W., Vermeer L., Dmitriev P. “Diagnosing Sample Ratio Mismatch in Online Controlled Experiments: A Taxonomy and Rules of Thumb for Practitioners.” KDD 2019 — the p < 0.0005 convention and the taxonomy of causes used in this page.
    2. Kohavi R., Tang D., Xu Y. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press, 2020 — SRM as a trust guardrail, and why a mismatched experiment should not be analysed.
    3. Cochran W. G. “The χ² test of goodness of fit.” Annals of Mathematical Statistics 23(3), 1952 — the expected-count rule of thumb behind the warning on this page.
    4. The chi-square upper tail is computed as the regularised incomplete gamma function (series and continued-fraction expansions), so the p-value is exact for any number of variants rather than a normal approximation.

    FAQ

    What is a sample ratio mismatch?

    A statistically implausible difference between the traffic split an experiment was configured with and the split actually observed in the data. It signals that something in assignment, logging, filtering or the analytics join is wrong, which makes the experiment’s metrics unreadable until the cause is found.

    What p-value means there is an SRM?

    The convention among experimentation platforms is p < 0.0005 from a chi-square goodness-of-fit test. It is far stricter than 0.05 because the check runs on every experiment, often daily, so a 5% false-alarm rate would produce constant noise and get ignored.

    Can I check SRM with more than two variants?

    Yes. The chi-square test takes k − 1 degrees of freedom for k variants, which is what this page computes. Read the residual column to see which variant is responsible: each residual squared is that variant’s share of the total chi-square.

    My split was 90/10 — does that count as a mismatch?

    No. A mismatch is a deviation from the split you configured, not from an even split. Enter 90 and 10 as the configured shares and the test compares against those. A tool that assumes 50/50 will flag every ramped experiment.

    The check passed. Is my experiment trustworthy?

    It passed one check. SRM cannot see a bug that affects both variants equally, a metric defined differently per variant, or non-random bucketing that happens to be even. And on small samples a pass only rules out large imbalances — the detectable-imbalance figure says how large.

    Can I just drop the extra users and carry on?

    No. The users missing from one variant are not a random sample, so reweighting or trimming hides the bias instead of removing it. Find the mechanism, fix it and rerun the experiment.

    Should I check SRM every day?

    Checking daily is normal practice, and it is why the threshold is 0.0005 rather than 0.05. Look for a mismatch that persists or grows across days rather than acting on a single flag, and remember that checking many segments brings back the multiple-comparison problem.

    More in A/B testing

    Open the collection →

    Tours, tooltips and checklists whose impact shows up in the numbers.

    Try Clew for freeHow it worksFree plan forever · no credit card

    Product tours, popups and onboarding in the age of AI

    7 patterns with step counts, copy rules and what to measure, what AI changes, and a checklist before you publish. PDF, 2 pages. What is inside →