Premium
Scoring Stability in a Large‐Scale Assessment Program: A Longitudinal Analysis of Leniency/Severity Effects
Author(s) -
Palermo Corey,
Bunch Michael B.,
Ridge Kirk
Publication year - 2019
Publication title -
journal of educational measurement
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 1.917
H-Index - 47
eISSN - 1745-3984
pISSN - 0022-0655
DOI - 10.1111/jedm.12228
Subject(s) - summative assessment , psychology , scale (ratio) , rating scale , stability (learning theory) , longitudinal study , trait , multilevel model , clinical psychology , statistics , developmental psychology , formative assessment , mathematics education , mathematics , computer science , physics , quantum mechanics , machine learning , programming language
Although much attention has been given to rater effects in rater‐mediated assessment contexts, little research has examined the overall stability of leniency and severity effects over time. This study examined longitudinal scoring data collected during three consecutive administrations of a large‐scale, multi‐state summative assessment program. Multilevel models were used to assess the overall extent of rater leniency/severity during scoring and examine the extent to which leniency/severity effects were stable across the three administrations. Model results were then applied to scaled scores to estimate the impact of the stability of leniency/severity effects on students’ scores. Results showed relative scoring stability across administrations in mathematics. In English language arts, short constructed response items showed evidence of slightly increasing severity across administrations, while essays showed mixed results: evidence of both slightly increasing severity and moderately increasing leniency over time, depending on trait. However, when model results were applied to scaled scores, results revealed rater effects had minimal impact on students’ scores.