Clew

HomeToolsUX researchNASA-TLX calculator

NASA-TLX calculator: raw and weighted task load index

Score the NASA Task Load Index without the spreadsheet. Fill in one questionnaire — six ratings and, if you ran them, the 15 pairwise comparisons — or paste a whole study and get the group’s raw TLX and weighted TLX with a confidence interval, the workload profile across the six dimensions, and which dimension is actually driving the score.

● Free, no sign-upUpdated:

Six ratings (0–100) in questionnaire order: mental, physical, temporal, performance, effort, frustration. Add either six weight columns (each 0–5, summing to 15) or the 15 pairwise choices written as 1 or 2. Leading id columns and a header row are ignored.
Confidence level

—

 

Raw TLX— 
Weighted TLX— 
Confidence interval— 
Biggest driver— 
Weighting effect— 
Spread— 
Workload profile across the six dimensions
Workload profile across the six dimensions
mean ratingbiggest drivermean weight (0–5, on the same box)
Raw TLX per respondent
Raw TLX per respondent
respondentsmean with interval

Per dimension

Mean rating, standard deviation, mean weight and contribution to the weighted score for each of the six dimensions
DimensionMeanSD WeightContribution

Method. Raw TLX is the plain mean of the six ratings. Weighted TLX multiplies each rating by how often that dimension won a pairwise comparison and divides by 15. Contribution is that dimension’s share of the final score, so the six contributions add up to it. Confidence intervals use the Student t distribution on the respondent means, which assumes respondents are independent — fine for a between-subjects study, wrong if the same person rated several tasks.

Everything is calculated in your browser — nothing you enter is sent anywhere.

How to use it

  1. Run the task first, then the questionnaireTLX is retrospective: ask immediately after the task, about that task. Asking about “the product” instead of a task is the fastest way to get numbers that mean nothing.
  2. Decide about the weights before you collectThe 15 pairwise comparisons take a few minutes per respondent. Raw TLX without them correlates closely with the weighted score in most studies, so most teams skip them — but decide once, for everyone, not per respondent.
  3. Check which way Performance pointsOn the original form Performance runs Perfect → Failure, so a high number means poor performance and high workload. If your survey tool recorded it the other way round, tick the box on this page.
  4. Paste the study, not the averagesOne row per respondent. Averaging before you paste loses the spread, the confidence interval and the per-dimension standard deviations — which is where the finding usually is.
  5. Read the profile, not just the scoreTwo tasks can both score 60 with completely different profiles: one all mental demand, one all frustration. The table and the bars tell you which to fix.

How NASA-TLX is scored

The NASA Task Load Index was developed by Sandra Hart and Lowell Staveland at NASA Ames in the 1980s and has since been used in several thousand studies. It measures perceived workload on six dimensions, each rated 0–100 on a 21-point line:

mental demand how much thinking, deciding, searching, remembering physical demand how much bodily activity temporal demand how much time pressure performance how successful you were (Perfect 0 → Failure 100) effort how hard you had to work for that performance frustration how insecure, discouraged, irritated or annoyed you felt raw TLX (RTLX) = (Σ ratings) / 6 weighted TLX = Σ (rating_i · weight_i) / 15 weight_i = how many of the 15 pairwise comparisons dimension i won (0…5)

The original procedure adds a weighting step: the respondent is shown all 15 pairs of dimensions and picks, for each pair, the one that contributed more to the workload of this task. A dimension can win at most five comparisons, and the weights always sum to 15. The idea is that time pressure matters more in air-traffic control than in proofreading, and the respondent should say so rather than the analyst assuming it.

Raw TLX, and why almost everybody uses it

Hart’s own 2006 review of 20 years of TLX use found that most researchers had quietly dropped the weighting step and averaged the six ratings instead, and that this “raw TLX” behaves about as well — sometimes more sensitively, sometimes less, rarely differently enough to change a conclusion. That is why this page shows both and reports the gap between them: if the weighting moves your score by a point, it did not matter; if it moves it by ten, your respondents are telling you something about which dimension dominates.

Performance is the reversed one

Five dimensions run low → high in the direction of more workload. Performance runs Perfect → Failure, so on the original form a high number means the person did badly, which counts as high workload. Survey tools routinely record it the friendly way round (good = 100), and the resulting scores are wrong in a way that is hard to spot because they still look plausible. The checkbox on this page flips it.

When NASA-TLX misleads you

TLX is a well-validated instrument that is easy to misuse. These are the failure modes that matter.

  • There is no universal “good” score. This page deliberately does not print a grade, because a global benchmark for TLX does not exist in the way it does for the System Usability Scale. Published scores vary enormously by domain and by task — a demanding cockpit task and a demanding spreadsheet task are not on a shared scale. Grier’s 2015 meta-analysis of around a thousand tasks is the place to look for distributions by task type, and its central message is that comparisons are only meaningful within a task type. Any tool that hands you an A–F grade for a TLX score has invented it.
  • It is only useful as a comparison. TLX earns its keep comparing two designs, two tasks or the same task before and after a change — with the same participants, the same task wording and the same moment of asking. A single number for a single condition is almost uninterpretable.
  • It measures perceived workload, not performance or usability. A task can feel easy and be done wrong, or feel hard and be done perfectly. TLX does not replace task success, time on task or the System Usability Scale; it explains part of why they came out as they did.
  • The Performance subscale confuses respondents as well as analysts. Some people rate how well they think they did; some rate how satisfied they are; some read the reversed anchors the wrong way. If Performance is your biggest driver, look at the raw answers before believing it.
  • Averaging the six dimensions hides the finding. The score is a sum of things that are not commensurable. Two conditions with the same mean can differ completely in profile, and the profile is what you can act on.
  • Weights are about the task, not the person. The pairwise comparisons must be answered for the specific task just performed. Reusing one respondent’s weights across several tasks, or averaging weights and applying them to everyone, quietly turns the weighted score back into a raw one with extra steps.
  • Repeated measures break the confidence interval. The interval on this page treats each row as an independent respondent. If the same people rated several tasks, you need a within-subjects analysis; the interval shown here will be too narrow.
  • Retrospective ratings drift. Asked ten minutes later, people rate the peak of the task rather than its average. Ask immediately, and ask the same way every time.
  • Small samples are normal and still small. Eight participants is a common TLX study and a wide interval. Report the interval and the profile; do not report a two-point difference between conditions as a result.

Used as a paired comparison with a stable protocol, TLX is one of the most useful instruments in the toolbox — it tells you which kind of hard a design is. Used as an absolute score with a made-up grade, it is decoration.

Using TLX on a software product

Most TLX literature is about aircraft, operating theatres and driving simulators, but the instrument transfers to software work with two adjustments.

  • Pick a task, not a screen. “Invite a teammate and give them edit access”, not “the settings page”. Workload is a property of what someone is trying to do.
  • Expect physical demand to be near zero and keep it anyway. It is a useful floor: a respondent rating physical demand at 60 for a web task has misread the scale, and the row is worth checking.
  • First-run tasks are the interesting ones. Mental demand and frustration during a first attempt are what onboarding is supposed to reduce, and they move faster than satisfaction scores do. Pair TLX with time to value and the funnel to see whether a lighter-feeling flow actually converts.
  • Compare before and after the same way. Same task text, same point in the session, same participants if you can get them. Then a five-point drop in mental demand is evidence rather than noise.
  • Watch the tooltip and tour copy. Temporal demand and frustration climb when in-product guidance moves faster than people read. The tooltip reading time calculator and the tooltip copy checker deal with exactly that.

Sources and licence

  1. Hart S. G., Staveland L. E. “Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research.” In Hancock P. A., Meshkati N. (eds), Human Mental Workload, North-Holland, 1988 — the instrument, the six dimensions and the pairwise weighting procedure.
  2. Hart S. G. “NASA-Task Load Index (NASA-TLX); 20 Years Later.” Proceedings of the Human Factors and Ergonomics Society Annual Meeting, 2006 — how TLX is used in practice, and the standing of raw TLX.
  3. Grier R. A. “How High is High? A Meta-Analysis of NASA-TLX Global Workload Scores.” Proceedings of the Human Factors and Ergonomics Society Annual Meeting, 2015 — distributions of global TLX by task type; the reason this page does not print a universal grade.
  4. NASA Human Systems Integration Division, NASA Task Load Index: NASA states that TLX is open source, available for use worldwide, and that permission is not required either to use it or to modify it. This page implements the published scoring and is not affiliated with or endorsed by NASA.

FAQ

What is NASA-TLX?

A questionnaire that measures perceived workload on six dimensions — mental demand, physical demand, temporal demand, performance, effort and frustration — each rated 0 to 100 immediately after a task. It was developed at NASA Ames in the 1980s and is used in several thousand published studies.

How do you calculate the NASA-TLX score?

Raw TLX is the average of the six ratings. Weighted TLX multiplies each rating by how often that dimension was chosen in the 15 pairwise comparisons and divides the total by 15. This page computes both, and the difference between them.

What is the difference between raw TLX and weighted TLX?

Weighted TLX uses each respondent’s own judgement of which dimensions mattered for this task; raw TLX treats all six as equally important. Most modern studies use raw TLX, which Hart’s 2006 review found behaves about as well. A large gap between the two tells you the workload is concentrated in a few dimensions.

What is a good NASA-TLX score?

There is no universal answer, and this page will not invent one. TLX scores vary widely by domain and by task, so a score is only interpretable against another score for a comparable task — the same task before and after a change, or two designs tested the same way. Grier’s 2015 meta-analysis gives distributions by task type if you need external reference points.

Do I have to do the 15 pairwise comparisons?

No. Raw TLX skips them and is what most studies now report. Do them if you expect the dimensions to matter very unequally for your task, and if you do, collect them from every respondent — mixing weighted and unweighted rows means the two scores describe different groups, which this page warns about.

Why is the Performance scale reversed?

On the original form Performance runs from Perfect to Failure, so a high number means the person performed poorly, which counts as high workload. Survey tools often record it the other way round. Tick the checkbox on this page to flip the column, and check the score moves the way you expected.

How many participants do I need?

TLX studies are often run with 8 to 20 participants, which gives a wide confidence interval. There is no magic number: report the interval rather than the mean alone, and be careful about calling small differences between conditions a result.

Is my data uploaded anywhere?

No. Parsing, scoring and the charts all run in your browser. Only what you put into the URL yourself, by copying the share link, leaves the page.

More in UX research

Open the collection →

Tours, tooltips and checklists whose impact shows up in the numbers.

Try Clew for freeHow it worksFree plan forever · no credit card

Product tours, popups and onboarding in the age of AI

7 patterns with step counts, copy rules and what to measure, what AI changes, and a checklist before you publish. PDF, 2 pages. What is inside →