Clew

HomeToolsProduct managementCSAT & CES calculator

CSAT & CES calculator with period comparison

Enter the distribution of satisfaction ratings (CSAT, 1–5) or effort ratings (CES, 1–7) and the calculator returns the share of satisfied customers, the mean rating and their confidence intervals. Add the previous period and it tests whether the metrics really changed or it is just noise.

● Free, no sign-upUpdated:

Question: “How satisfied were you with…?” from 1 (very dissatisfied) to 5 (very satisfied). Enter how many times each rating was given.

Responses per rating

NowBefore1 — very dissatisfied2345 — very satisfied
Fill from a list of ratings
Numbers from 1 to 5 separated by spaces, commas or new lines — for example, a column from a survey export.
Confidence level

—

 

CSAT now— 
CSAT before— 
Share change— 
Mean now— 
Mean before— 
Mean change— 
CSAT rating distribution, share of responseshover a rating
CSAT rating distribution, share of responses
NowBeforeratings 4–5 highlighted

Formulas. CSAT = (4s + 5s) / all responses; Wilson interval. Period difference: two-proportion z-test; difference in means: Welch’s t-test.

Everything is calculated in your browser — nothing you enter is sent anywhere.

How to use it

  1. Pick a metricCSAT is satisfaction with a specific interaction on a 1–5 scale. CES is how easy it was to get a task done, on a 1–7 agreement scale.
  2. Enter the distributionHow many times each rating was given — from your survey tool’s report. Or paste a list of ratings and let the calculator count them.
  3. Add the previous periodThe same question last month, before a release or in a control group. The periods must be independent samples.
  4. Read the verdictThe calculator compares both the satisfied share and the mean rating, and shows intervals for the difference and p-values. If an interval includes zero, the change is not proven.

How to calculate CSAT: top-2-box and mean rating

CSAT (Customer Satisfaction Score) measures satisfaction with a specific moment: a support conversation, a purchase, finishing setup. The most common version is the share of satisfied customers — ratings of 4 and 5 on a five-point scale, known as “top-2-box”:

CSAT = (number of 4s + number of 5s) / all responses × 100% Example: 6, 9, 22, 88 and 125 responses for ratings 1…5, 250 in total CSAT = (88 + 125) / 250 = 85.2% 95% Wilson interval: 80.3% … 89.1%

The mean rating (“4.27 out of 5”) is useful too, but very different distributions can share the same mean. That is why the calculator shows both numbers and the full distribution. The share uses a Wilson interval: at shares around 80–90% the textbook “share ± margin” formula produces noticeably skewed bounds.

If you need an overall loyalty index rather than a rating of a single interaction, calculate NPS with the NPS calculator, which also gives the margin of error and a period comparison.

How to calculate Customer Effort Score

CES (Customer Effort Score) shows how easy it is for a customer to get something done. The metric followed Dixon, Freeman and Toman’s 2010 Harvard Business Review article, which argued that loyalty depends more on reducing effort than on trying to delight. The current version asks for agreement with “The company made it easy for me to handle my issue” on a 1–7 scale:

CES = mean rating Interval: CES ± t(n − 1) · s / √n “Easy” share = ratings 5–7 / all responses Example: 160 responses, mean 5.34, s = 1.40 95% interval: 5.12 … 5.56

Beyond the mean, look at the share of 1–3 ratings: these people found it hard, and their journeys are the first to investigate. CES is especially useful after onboarding, data import or checkout — wherever the interface gets in the way most.

Did CSAT or CES change significantly?

When CSAT goes from 78% to 85%, the first question is whether it is chance. The calculator tests two hypotheses independently.

Difference in shares: z-test

p̂ = (x₁ + x₂) / (n₁ + n₂) pooled share z = (p₂ − p₁) / √( p̂ (1 − p̂) (1/n₁ + 1/n₂) ) Interval for the difference: (p₂ − p₁) ± z_crit · √( p₁(1−p₁)/n₁ + p₂(1−p₂)/n₂ ) Example: before 195 of 250 (78.0%), now 213 of 250 (85.2%) Difference +7.2 pp, z = 2.08, p-value = 0.038, interval +0.4 … +14.0 pp

Difference in means: Welch’s t-test

t = (x̄₂ − x̄₁) / √( s₁²/n₁ + s₂²/n₂ ) df = (s₁²/n₁ + s₂²/n₂)² / ( (s₁²/n₁)²/(n₁−1) + (s₂²/n₂)²/(n₂−1) )

Welch’s test does not assume equal variances and works with samples of different sizes. Ratings are an ordinal scale, so strictly speaking a mean is a convention; with hundreds of responses the t-test is reliable, but a 0.05-point difference is not worth interpreting even at p < 0.05.

Limitations

  • The periods must be independent: different respondents or different conversations. If the same people answered twice, you need a paired test.
  • The test covers random error only. If the new survey was shown in a different place or to a different audience, the difference may come from the method rather than the product.
  • Do not re-check significance every day until it “works”: repeated peeking guarantees false positives. The survey sample size calculator tells you how many responses to plan for.

Sources

  1. Dixon M., Freeman K., Toman N. Stop Trying to Delight Your Customers. Harvard Business Review, July–August 2010 — Customer Effort Score.
  2. Dixon M., Toman N., DeLisi R. The Effortless Experience. Portfolio/Penguin, 2013 — CES on a 1–7 scale.
  3. Welch B. L. The Generalization of “Student’s” Problem when Several Different Population Variances are Involved. Biometrika, 34(1/2), 1947 — t-test for unequal variances.
  4. Wilson E. B. Probable Inference, the Law of Succession, and Statistical Inference. Journal of the American Statistical Association, 22(158), 1927.
  5. ISO 10004:2018. Quality management — Customer satisfaction — Guidelines for monitoring and measuring.

FAQ

How do you calculate CSAT?

Divide the number of 4 and 5 ratings by all responses on a 1–5 scale and multiply by 100%. For example, 213 satisfied out of 250 responses is a CSAT of 85.2%.

What is top-2-box?

The share of responses in the two highest options of a scale: for CSAT on a 1–5 scale, the 4s and 5s. It is easier to interpret than a mean and more robust to different distributions.

How do you calculate CES?

In the 1–7 version, CES is the mean agreement with “It was easy to get my issue resolved”. It also helps to track the share of 5–7 ratings (“easy”) and 1–3 ratings (“difficult”).

How do I know if CSAT changed significantly?

Compare the two periods’ shares with a z-test: if the p-value is below 0.05 (at 95% confidence) and the interval for the difference excludes zero, the change is significant. The calculator does this automatically.

Which test should I use for a difference in mean ratings?

Welch’s t-test: it does not assume equal variances and suits samples of different sizes. For ordinal ratings with few responses, check the result against the distribution as well.

How are CSAT, CES and NPS different?

CSAT is satisfaction with a specific interaction, CES is how easy a task was, NPS is willingness to recommend the product overall. CSAT and CES are collected right after an event, NPS less often, typically quarterly.

How many responses do I need to compare periods?

It depends on the difference you want to detect. With 80% power at a 5% significance level, detecting a rise in CSAT from 80% to 85% needs about 900 responses per period; from 80% to 90%, about 200.

More in Product management

Open the collection →

Tours, tooltips and checklists whose impact shows up in the numbers.

Try Clew for freeHow it worksFree plan forever · no credit card

Product tours, popups and onboarding in the age of AI

7 patterns with step counts, copy rules and what to measure, what AI changes, and a checklist before you publish. PDF, 2 pages. What is inside →