HealthMeasures Scores Are Computed as T-Scores
HealthMeasures were developed with item response theory (IRT) to produce precise, powerful scores that are comparable across forms. These scores are calculated as T-scores using multiple tools.
Scoring Options
You can calculate a HealthMeasures score in three ways. API-powered scoring and the HealthMeasures Scoring Service are preferred over manual scoring.
API-Powered Scoring
The Assessment Center API is integrated software that powers the Epic PROMIS® Application, REDCap, and many other platforms to administer and score HealthMeasures. It can also be installed on your desired server to generate scores locally.
HealthMeasures Scoring Service
The HealthMeasures Scoring Service is a free, web-based application that scores a data file of raw participant responses and returns by email a file with calculated T-scores for all measures.
Manual Scoring
If API-powered scoring and the HealthMeasures Scoring Service are not options, you can calculate a score by following the instructions in the relevant scoring manual.
More About Scoring
Preference-Based Scores
Preference-based scores provide an overall summary of health-related quality of life on a common 0 to 1 metric.
Report Scores
Report HealthMeasures T-scores and Standard Errors as integers. Follow best practices in displaying results visually.
Score Translations
Translated HealthMeasures use the same scoring as the English measures. Follow the scoring instructions described above.
Frequently Asked Questions About Scoring
If you have missing data, use the HealthMeasures Scoring Service or API-powered scoring. Both of these approaches calculate scores based on individual item responses (“response pattern scoring”) rather than summed totals. This means that any subset of items can be used to produce a score and a valid score can be produced even when there is missing data. However, too much missing data reduces the precision of an individual score. If you need a precise score, make sure you have responses to at least 4 items.
Manual scoring is likely to produce a slightly different T-score than the HealthMeasures Scoring Service and API-powered scoring approaches. This is because the HealthMeasures Scoring Service and API-powered scoring use a more precise scoring methodology (“response pattern scoring”) that takes into account information about each response option to every item. Manual scoring uses simpler tables that convert a raw summed score to a T-score. The challenge with the summed score approach is that there are many ways (i.e., different response patterns) to get to the same raw summed score. For example, responses 1, 1, 2, 2 and responses 3, 1, 1, 1 both yield a sum of 6. The summed score approach uses weighted probabilities for each response pattern to calculate the most likely T-score associated with a raw summed score. The software-based scoring approaches do not need to make this estimation and calculate a separate T-score for each response pattern. As a result, there is less precision (more error) with manual scoring. This can make more of a difference for interpreting individual scores or score means for small groups than it does for interpreting score means for large groups. If you can use response pattern scoring, do it!
The HealthMeasures Scoring Service does not collect any protected health information from your respondents. Only the assessment number (e.g., 1, 2, 3), item IDs, raw score responses (e.g., 1, 2, 3, 4, 5), and participant identifying number (whatever alphanumeric code you wish) are included in the spreadsheet you upload. The Scoring Service does collect the email address where you would like to have the scored file sent.
If you would like to keep all data behind your firewall, use API-powered scoring.
Yes! First, we don’t recommend administering a full item bank. You are not gaining any benefit from administering substantially more items than you see in a short form. Second, any combination of items from the same item bank can produce a T-score. All T-scores from measures from the item bank are comparable. This means that scores from a 4-, 6-, 8-, 10-, and 20-item short form, computer adaptive test, or full item bank can be compared to each other.
Yes! Every item is calibrated individually. Every item could be administered alone and you can calculate a T-score from the raw response score. However, a single item alone produces a T-score with significantly more error. Administer at least 4 items if you need a precise score.
Yes! All items in a PROMIS item bank can have their raw response score transformed into a T-score. The T-score for a single item, a set of items in a short form, and a computer adaptive test are all directly comparable. All T-scores have a mean of 50 and SD of 10.
Only API-powered platforms can calculate scores for CATs because the measures require specialized software to select each item that is administered to a given individual. The software automatically produces the resulting T-score.
HealthMeasures scores have been developed by domain and are not intended to be combined across domains. However, preference-based scores can be calculated if the goal is to evaluate overall quality of life for cost-effectiveness or utility analyses. Preference-based scores produce a single numerical index across domains; however, they are not appropriate for evaluating individuals.
Refer to Score Cut Points for more information. All of the score interpretation figures show a range of 20 to 80 T-score points. This reflects the mean (T=50) +/- 3 standard deviations (SD=10). Almost everyone (>99%) will fall within this range based on the normal distribution.
However, the actual range of possible scores varies by measure. Some measures include items that assess different levels of the concept at more extreme levels, allowing for scores outside the 20-80 range (e.g., as low as T=9 or as high as T=85), even though few people will receive such extreme scores.
The PROMIS-29 is not designed to produce an overall score. Rather, the individual scores on each of the 7 domains are intended to provide a comprehensive summary of function across mental health, physical health, and social health. However, the PROMIS-29 can allow for computation of a PROPr score, a PROMIS preference-based score. Learn more about PROPr in Preference-Based Scores.
This means when the measure was developed, the item statistics did not meaningfully differentiate between the two response options. These are called “collapsed categories” as the response options are intended to be collapsed together for scoring.
Some HealthMeasures (e.g., PROMIS Physical Function) have multiple versions (v1.0, v1.1, v1.2, v2.0, etc.). Sometimes the scores on different versions of measures can be compared to each other. Sometimes they can’t be compared. To identify what is comparable, view Compare Versions.
Last updated: September 14, 2026