State of the psychometric methods: patient-reported outcome measure development and refinement using item response theory

Angela M. Stover; Lori D. McLeod; Michelle M. Langer; Wen-Hung Chen; Bryce B. Reeve

doi:10.1186/s41687-019-0130-5

Journal of Patient-Reported Outcomes (Jul 2019)

State of the psychometric methods: patient-reported outcome measure development and refinement using item response theory

Angela M. Stover,
Lori D. McLeod,
Michelle M. Langer,
Wen-Hung Chen,
Bryce B. Reeve

Affiliations

Angela M. Stover: Department of Health Policy and Management, University of North Carolina at Chapel Hill
Lori D. McLeod: RTI Health Solutions
Michelle M. Langer: Lineberger Comprehensive Cancer Center, University of North Carolina at Chapel Hill School of Medicine
Wen-Hung Chen: RTI Health Solutions
Bryce B. Reeve: Department of Health Policy and Management, University of North Carolina at Chapel Hill

DOI: https://doi.org/10.1186/s41687-019-0130-5
Journal volume & issue: Vol. 3, no. 1
pp. 1 – 16

Abstract

Read online

Abstract Background This paper is part of a series comparing different psychometric approaches to evaluate patient-reported outcome (PRO) measures using the same items and dataset. We provide an overview and example application to demonstrate 1) using item response theory (IRT) to identify poor and well performing items; 2) testing if items perform differently based on demographic characteristics (differential item functioning, DIF); and 3) balancing IRT and content validity considerations to select items for short forms. Methods Model fit, local dependence, and DIF were examined for 51 items initially considered for the Patient-Reported Outcomes Measurement Information System® (PROMIS®) Depression item bank. Samejima’s graded response model was used to examine how well each item measured severity levels of depression and how well it distinguished between individuals with high and low levels of depression. Two short forms were constructed based on psychometric properties and consensus discussions with instrument developers, including psychometricians and content experts. Calibrations presented here are for didactic purposes and are not intended to replace official PROMIS parameters or to be used for research. Results Of the 51 depression items, 14 exhibited local dependence, 3 exhibited DIF for gender, and 9 exhibited misfit, and these items were removed from consideration for short forms. Short form 1 prioritized content, and thus items were chosen to meet DSM-V criteria rather than being discarded for lower discrimination parameters. Short form 2 prioritized well performing items, and thus fewer DSM-V criteria were satisfied. Short forms 1–2 performed similarly for model fit statistics, but short form 2 provided greater item precision. Conclusions IRT is a family of flexible models providing item- and scale-level information, making it a powerful tool for scale construction and refinement. Strengths of IRT models include placing respondents and items on the same metric, testing DIF across demographic or clinical subgroups, and facilitating creation of targeted short forms. Limitations include large sample sizes to obtain stable item parameters, and necessary familiarity with measurement methods to interpret results. Combining psychometric data with stakeholder input (including people with lived experiences of the health condition and clinicians) is highly recommended for scale development and evaluation.

Published in Journal of Patient-Reported Outcomes

ISSN: 2509-8020 (Online)
Publisher: SpringerOpen
Country of publisher: United Kingdom
LCC subjects: Medicine: Public aspects of medicine
Website: https://jpro.springeropen.com/

About the journal

Abstract

Keywords