Score Cut Points
Definition
Score cut points assign descriptive words (e.g., moderate) to a score range (e.g., 60–70). These words can make interpreting a score more intuitive and actionable. They can also help you quickly understand how a respondent is doing, identify changes over time, and prioritize areas for attention. Score cut point figures show score ranges and descriptive labels.
How to Use Score Cut Point Figures
- Identify your respondent (adult, pediatric, parent proxy, or early childhood parent report).
- Identify your measure’s domain (e.g., Depression, Fatigue, Physical Function).
- Select the correct figure for your measure and respondent from the figures on the measure system’s Score Cut Point page.
Tips
- Cut points (also known as thresholds) vary by domain. On the Score Cut Point pages, each domain is shown in a figure with the appropriate descriptive terms and cut points.
- PROMIS® domains use different terms to describe score ranges. For example, some domains use “within normal limits/mild/moderate/severe” whereas others use “very high/high/average/low/very low.”
- The same cut points are used for short forms, computer adaptive tests (CATs), scales, and profiles for the same domain.
- Sometimes score cut points vary by the version of a measure. For example, some PROMIS Pediatric measures have different score cut points for v2 compared to v3.
- Note that although the figures range from about 20–80, not all measures for all domains can generate scores across that full range.
- Score cut points were usually identified by reviewing the distribution of scores. This is a norm-based method. Learn more about methods for selecting score cut points.
Example
In the figure below, scores less than or equal to 55 T-score points are described as “within normal limits.” Scores between 55 and 60 are described as “mild” whereas scores 60–70 are labeled “moderate” and scores above 70 are labeled “severe.” The figure also shows that the general population has a mean of 50 and standard deviation of 10. Therefore, if a patient has a score of 60, you know that the patient is one standard deviation worse than the general population average and on the border between mild and moderate.
Score Cut Points Vary Across Measures
Score cut points can be used with any measure (e.g., short form, computer adaptive test (CAT), profile). Find the score cut points for the measure you need by first identifying the respondent.
Adult
There are multiple score cut points for adult HealthMeasures.
Pediatric and Parent Proxy
Early Childhood Parent-Report
Score Ranges
All of the score interpretation figures show a range of 20 to 80 T-score points. This reflects the mean (T=50) +/- 3 standard deviations (1 SD = 10). Almost everyone (>99%) will have a score in this range based on the normal distribution.
However, the actual range of possible scores varies by measure. Some measures include items that assess different levels of the domain at more extreme levels. This allows for scores outside the 20-80 range. In some cases, this is as low as T=9 and as high as T=85. Note that few people will receive such extreme scores.
Find Score Ranges for Your Measure
To identify the exact score range for your measure, do the following:
For short forms, consult the scoring tables in the appropriate scoring manual. Each table shows the range of possible scores for that short form. For example:
- PROMIS Short Form v1.0 – Anxiety 4a: T=40.3 to T=81.6
- PROMIS Short Form v1.0 – Anxiety 8a: T=37.1 to T=83.1
For computer adaptive tests (CATs) there are no scoring tables. To estimate the possible score range, enter test responses in your data collection platform:
- In a CAT, select the response options with the highest scores for all items to estimate the upper limit.
- In a CAT, select the response options with the lowest scores for all items to estimate the lower limit.
Standard Setting Using Bookmarking
The HealthMeasures score cut point figures were established using norm-based methods. This means that researchers selected cut points by reviewing the distribution of scores.
Bookmarking is a different method for establishing score cut points. This method has been commonly used in educational testing to identify thresholds for levels of academic outcomes (e.g., math proficiency levels). It has also been used to establish thresholds for severity levels (e.g., no problems, mild problems, moderate problems, severe problems) for patient-reported outcome measures in multiple patient populations.
Bookmarking-based score thresholds exist for:
- PROMIS Pain Interference, Fatigue, Anxiety, and Depression in cancer (Cella et al., 2014)
- PROMIS Physical Function, Cognitive Function, and Sleep Disturbance in cancer (Rothrock et al., 2019)
- PROMIS Pain Interference and Fatigue in rheumatoid arthritis (Bingham et al., 2021)
- PROMIS Physical Function, Pain Interference, Sleep Disturbance, and Depression in rheumatic conditions (Nagaraja et al., 2018)
- PROMIS Pain Interference in people living with chronic pain (Cook et al., 2025)
- PROMIS Physical Function, Upper Extremity Function, and Pain Interference in patients with bone fractures (Rothrock et al., 2023; Rothrock et al., 2025)
- PROMIS Pediatric Mobility, Upper Extremity Function, Pain Interference, and Fatigue in juvenile idiopathic arthritis (Morgan et al., 2017)
- PROMIS Pediatric Anxiety, Depressive Symptoms, Fatigue, and Mobility in juvenile idiopathic arthritis and childhood-onset systemic lupus erythematosus (Mann et al., 2020)
- Neuro-QoL Cognitive Function and Ability to Participate in Social Roles and Activities in adults with acquired cognitive and language disorders (Cohen et al., 2023)
- Neuro-QoL fatigue, physical function, and sleep disturbance in multiple sclerosis (Cook at al., 2015)
Bolstering Your Evidence for Using a Specific Cut Point
- The recommended score cut point figures on HealthMeasures.net are a great starting point for understanding the meaning of a given score. However, there is no single number that reflects the exact transition point between levels of severity for all patients in all circumstances for all purposes.
- These recommended cut points serve as general guidelines. There are different research methods for setting cut points, and they can yield different values. Also, analyses in different samples and contexts are likely to lead to different values. Therefore, selecting a cut point requires a judgment on your part.
- Assembling multiple sources of evidence can make a stronger argument for using a given cut point. The HealthMeasures score interpretation cut point figures are one source and are fine to use on their own. Publications using standard setting methods (e.g., bookmarking) as well as those reporting descriptive statistics on your target patient population are additional sources.
- Consider how you will use the cut point. For example, for a score at an identified cut point (e.g., T=60), you should consider the consequences of assigning it to the less severe (e.g., mild) or more severe (e.g., moderate) category. Evaluate the relative risk of describing someone with moderate symptoms as mild (similar to a false positive) versus the risk of describing someone with mild symptoms as moderate.
- Calculating the cumulative percentile associated with different T-scores can also help you gauge how many people within the general population would be at or below a given T-score. See How to Calculate Percentiles for instructions.
Other Methods
Patient Acceptable Symptom State (PASS)
The PASS measure uses a person-centered “Yes/No” question to indicate whether a patient is satisfied with their health given their personal activity levels, pain, and functional impairment. PASS can be useful to guide provider-patient discussions and determine patient priorities.
Some research indicates that individual differences in patients affect the threshold for achieving PASS. For example, studies using PROMIS Physical Function have found that:
- Patients who made less than $25,000 USD per year needed a PROMIS score five points lower to achieve PASS compared to patients who made more than $100,000 USD per year (Bernstein et al., 2019).
- Women required four fewer points on PROMIS to achieve PASS compared to men (Baumhauer et al., 2018).
- Athletes needed a PROMIS score about eight points higher than non-athletes to achieve PASS (Kuhns et al, 2020).
The following table presents guidelines based on specific populations because a global connection between PROMIS Physical Function and PASS has not been established.
PROMIS Physical Function Scores that Correspond to PASS Achievement
| Patient Population | PROMIS PF Score | Citation | Number of Patients Studied |
|---|---|---|---|
| Primary Care | |||
| Rural Outpatient Clinic | 43 | Jacobson et al., 2020 | 360 |
| Patients with Musculoskeletal Issues | 45 | Houck et al., 2021 | 94 |
| Orthopedics | |||
| Foot and Ankle | 45 | Bernstein et al., 2019 | 2597 |
| Foot and Ankle | 50 | Baumhauer et al., 2018 | 450 |
| Foot and Ankle | 45 | Anderson et al., 2018 | 88 |
| Knee Arthroscopy | 46 | Lu et al., 2020 | 60 |
| Hip Arthroscopy | 52 | Kuhns et al., 2020 | 113 |
| Hip Arthroscopy | 47 | Bodendorfer et al., 2021 | 124 |
| Disease-Specific | |||
| Systemic Lupus Erythematosus | 57 | Katz et al., 2020 | 481 |
Dichotomize Responses
Researchers have dichotomized responses to items and then used patterns of responses to determine no/mild, moderate, and severe score cut points.
- PROMIS Pediatric Pain Interference and Pain Behavior in children with sickle cell disease (Singh & Panepinto, 2019).
Receiver Operating Characteristics Curves
Researchers have used Receiver Operating Characteristics (ROC) curves to select score cut points.
- PROMIS Pediatric Asthma Impact in children and adolescents with asthma (Burbank et. al., 2024).
Score Cut Point Methods
Most HealthMeasures score cut points were selected using norm-based methods.
PROMIS scientists used norm-based methods. They constructed the interpretation of scores for Profile domains (Anxiety, Depression, Fatigue, Pain Interference, Physical Function, Sleep Disturbance, and Ability to Participate in Social Roles & Activities) by reviewing the data collected in the large-scale calibration testing data (Cella et al., 2010; Rothrock et al., 2010). This helped their understanding of the range of scores typically observed in the general population. Next, they evaluated whether 0.5, 1.0, and 2.0 standard deviations were reasonable thresholds to use across domains. This was done by looking at the percentage of participants from large-scale calibration testing that would then fit into each category and setting cut points using clinical judgment.
PROMIS scientists used norm-based methods.
- Some domains, like Pain Behavior, Anger, and Gastrointestinal Symptoms, followed the same approach and interpretation as the PROMIS Core domains. For domains (e.g., Smoking – Nicotine Dependence) that were calibrated on a clinical population (e.g., smokers), cut points were made at equal intervals of standard deviations (SDs), so at 1, 2, and 3 SDs below or above the average score of 50.
- The same approach was taken for domains (e.g., Meaning and Purpose, Companionship) where 50 represents an estimate of the mean of the general population, but no clear interpretation is available for a particular level. For example, a symptom like depression can be moderate or severe, but these labels make less sense with a concept like meaning and purpose.
- For Alcohol Use, a lower threshold of 55 (vs. 60) was chosen as a potential threshold for problematic drinking, given its alignment to thresholds on other screening measures, such as the AUDIT.
- For some positive health domains (e.g., Spirituality), measure developers set cut points at 0.5 above and below the mean of the general population.
PROMIS scientists used similar norm-based methods for the PROMIS Pediatric version 2 measures. PROMIS scientists constructed the interpretation of scores for Profile domains (Anxiety, Depression, Anger, Fatigue, Pain Interference, Physical Function, and Peer Relationships) by reviewing the data collected in the large-scale calibration testing data (Irwin et al., 2010). This helped their understanding of the range of scores typically observed in the general population. Next, they evaluated if 0.5, 1.0, and 2.0 standard deviations were reasonable thresholds to use across domains. This was done by looking at the percentage of participants from large-scale calibration testing that would then fit into each category. They set the PROMIS parent proxy version 2 measures’ thresholds to match those used for the Pediatric version 2 measures.
For PROMIS pediatric (GenPop) version 3 measures, PROMIS scientists started with the score cut points used for version 2. They compared score distributions between version 2 and version 3. Based on this comparison, they increased the cut points used in version 2 by 5 T-score points (i.e., 0.5 SD) in version 3 to account for the differences in reference norms between the two versions. They set the PROMIS parent proxy (GenPop) version 3 measures’ thresholds to match those used for the pediatric (GenPop) version 3 measures.
PROMIS scientists used norm-based methods. The PROMIS Early Childhood development team reviewed the score distributions from multiple sources, including the norming study as well as large national and international datasets that included the measures. Cut points for provisional score interpretation were created through a consensus process, with input from content and clinical experts. Additional clinical validation studies will help further refine these cut points.
PROMIS scientists used a different approach with the PROMIS Global Health 10-item scale, specifically the anchor-based method. Below we provide a detailed description of these methods.
To generate PROMIS Global score cut points, we applied the anchor-based method (Shi et al., 2019). The PROMIS Global Physical Health scale has four items measuring general physical health, everyday physical activities, pain, and fatigue. PROMIS Global Mental Health also has four items measuring general mental health, quality of life, social relationships, and emotional problems. The PROMIS Global Physical Health scale score is highly correlated with a single general health item (Global01), and the PROMIS Global Mental Health scale score is highly correlated with a single quality of life item (Global02). Both items have been used on many health surveys. For this reason, the five-category Global01 item (rated from excellent to poor) was used as an anchor item to estimate the scores of PROMIS Global physical health at each level of physical health, and the Global02 item was used to estimate PROMIS Global mental health scores at each level of mental health.
We analyzed the PROMIS 1 Wave 1 dataset (N = 21,133), collected from the U.S. general population (Cella 2015, Liu et al., 2010). The PROMIS Global Health scale includes Global01, PROMIS Global physical health items, and PROMIS Global mental health items (including Global02), which were administered to all the subjects. We first recoded the responses of global items according to the PROMIS Global Health v1.2 scoring manual published on HealthMeasures.net and scored global item responses using item response theory (IRT) Expected A Posteriori (EAP) scoring method, equivalent to those obtained via the HealthMeasures Scoring Service. To be consistent with previous work (Hays et al., 2015), we calculated the mean score on the Global Physical Health scale (according to the five categories of the Global 01 item) and the Global Mental Health scale (according to the five categories of the Global02 item) and then calculated the least-square means for both scales. After the mean scores of each health level for each scale were identified, we calculated the cut points of each scale by taking the midpoint of two adjacent mean scores and established four cut points corresponding to five levels ranging from excellent to poor for each scale.
The least-square means were calculated based on the five-category general health item for the Physical Health scale and the five-category quality of life item for the Mental Health scale. The cut points were calculated as the midpoints of two adjacent mean scores. For example, the least-square mean score on the Global Physical Health scale for the poor general health level is 32. The least-square mean score for the fair general health level is 39. Therefore, the midpoint of 36 was established as the cut point between the poor and fair category for the Global Physical Health scale. The score interpretation figures show our recommended cut points.
We repeated this analysis with PROMIS Scale v1.2 – Global Health Mental 2a and PROMIS Scale v1.2 – Global Health Physical 2a and found that the cut points were all within one T-score point different from the cut points for the PROMIS Global Health v1.2 (the 10-item measure). Hence, we suggest applying the same cut points to Global Health Mental 2a and Global Health Physical 2a as used for the longer measure.
NIH Toolbox Emotion measures use a single cut point approach based on one standard deviation from the mean (T-scores of 40 and 60; Babakanyan et al., 2018). The NIH Toolbox development team established cut points using:
- Statistical approach: One standard deviation from the mean (T-scores of 40 and 60)
- Rationale: Captures approximately 16% of the population at each tail
- Terminology: “Potentially problematic” (not clinical severity labels) due to absence of gold standard clinical validations
Importantly, because these are based on statistical distributions rather than clinical diagnoses, a single “potentially problematic” score should not be over-interpreted in isolation. In fact, research shows that over 60% of healthy adults have at least one such score when completing multiple Emotion Battery measures (Ingram & Karr, 2024). Therefore, interpretation should emphasize patterns across multiple scales, clinical context, and construct directionality.
Last updated: September 14, 2026