Manual Scoring
Manual scoring is an option when API-powered scoring and the HealthMeasures Scoring Service are not feasible. Unlike these preferred approaches to scoring, manual scoring uses a table that converts a raw summed score to a T-score. Follow the instructions on this page to manually score a HealthMeasure.
Summed Scoring
HealthMeasures are based on item response theory. This means that each item has statistical information available that can be used to calculate a score. The HealthMeasures Scoring Service and API-powered Scoring use a precise scoring methodology (“response pattern scoring”) that takes into account this information about each response option for every item. This can result in different T-scores by response pattern for the same summed score. For example, a response pattern of 1-2-3-2 and a response pattern of 3-1-2-2 both have a summed score of 8, but the responses to individual items are different. Response pattern scoring accounts for differences in symptom severity and functional status that are measured by different items. For example, selecting “almost always”, a raw response score of 4, to a question about running 10 miles would result in a higher score than selecting the same response to a question about walking one block since running 10 miles represents greater physical function than walking a block.
Compared to response pattern scoring, manual scoring uses simpler tables that convert a raw summed score to a T-score. The summed score approach uses weighted probabilities for each response pattern to calculate the most likely T-score associated with a raw summed score. The software-based scoring approaches do not need to make this estimation and calculate a separate T-score for each response pattern. As a result, there is less precision (more error) with manual scoring. This can make more of a difference for interpreting individual scores or score means for small groups than it does for interpreting score means for large groups.
When to Use Manual Scoring
- HealthMeasures are collected on paper and need to be scored immediately by the measure administrator.
- You have complete data.
- You have administered existing short forms (not custom forms or computer adaptive tests).
- You want to compare scores for large groups of respondents (more than 75 people).
Scoring Translations
If you are scoring a translated HealthMeasure:
- Identify the measure’s English-language measure name.
- Follow the instructions.
Manual Scoring Instructions
PROMIS®, Neuro-QoL™, NIH Toolbox®, and ASCQ-Me® measures generate T-scores. If you are not able to use a digital platform that automatically calculates scores or the HealthMeasures Scoring Service, you will need to calculate T-scores by hand.
Evaluate Missing Data
Confirm there are responses to all items. If any respondent skipped an item, remove them rather than imputing or prorating their score. Manual scoring cannot be used with missing data.
Calculate Raw Summed Score
Calculate a summed score by totaling the item responses together. Each item usually has five response options ranging in value from one to five. Consequently, a 4-item short form can have a total summed score ranging from 4 to 20.
Access the Scoring Manual
Use the following section to access the scoring manual for the measure you used.
Access the Scoring Table
In the Scoring Manual Appendix, locate the score conversion table for your measure. This table lists all possible raw summed scores and the corresponding T-score and standard error (SE).
Find Your T-Score
Find your raw summed score and record the corresponding T-score and SE. Use the T-score in all analyses, publications, or other work.
Scoring Manuals
PROMIS
Scoring Manuals are available for each PROMIS short form, scale, and profile measure. All measures for a domain are included in a manual.
Neuro-QoL
Neuro-QoL (which includes HDQLIFE and TBI-CareQOL) has one comprehensive scoring manual available to download.
NIH Toolbox
NIH Toolbox Emotion measures can be manually scored by following these scoring instructions.
Frequently Asked Questions
The manual scoring tables should only be used when you have complete data. If there is missing data, the HealthMeasures Scoring Service or API-Powered Scoring are better approaches. Because of how HealthMeasures are scored, imputing or prorating missing data (i.e., inserting an estimate where there is missing data) is not recommended when precise scores are needed.
Manual scoring is likely to produce a slightly different score than the HealthMeasures Scoring Service and API-powered scoring approaches. This is because the HealthMeasures Scoring Service and API-powered scoring use response pattern scoring that takes into account information about every response option to every item. Manual scoring, however, uses tables that convert a raw summed score to a T-score. The problem is that there are many ways to get to the same raw summed score. For example, responses 1, 1, 2, 2 and responses 3, 1, 1, 1 both yield a sum of 6. Consequently, HealthMeasures scientists estimated the most likely T-score associated with a raw summed score. There is less precision (more error) with manual scoring. If you can use response pattern scoring, do it!
Last updated: September 14, 2026