Author
Listed:
- Rudolf Debelak
- Matthias Ziegler
Abstract
Recent developments in the field of artificial intelligence and machine learning allow the wide application of large language models for the evaluation of written text and other non-numerical data. When applied in the context of psychological and educational assessments, such models can be used for assigning scores to essays and other types of responses. In contrast to classical tests, essays do not consist of test items, which leads to specific challenges in the evaluation of testing standards for scores obtained from AI models that differ from those observed for classical ability tests and personality questionnaires. To address these challenges, we discuss the evaluation of validity, fairness, and reliability for scores obtained from models of artificial intelligence in the context of automated essay scoring. We review existing methods, propose new methods, and further illustrate the reviewed methods with an empirical example based on the Hewlett Foundation data set on automated essay scoring. By applying the proposed framework to an evaluation based on a DistilBERT model, we find the model to be robust with sufficiently high internal consistency (Spearman-Brown coefficients in the range from .77 to .92). We further found empirical evidence for the validity of the evaluation model, but also indications for violations of fairness when comparing the human and AI scores across different topics. This study provides a standardized, replicable toolkit for researchers and practitioners to evaluate the psychometric quality of AI-based assessments.
Suggested Citation
Rudolf Debelak & Matthias Ziegler, 2026.
"Testing standards for AI-based scores in automated essay scoring,"
PLOS ONE, Public Library of Science, vol. 21(7), pages 1-29, July.
Handle:
RePEc:plo:pone00:0354680
DOI: 10.1371/journal.pone.0354680
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pone00:0354680. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosone (email available below). General contact details of provider: https://journals.plos.org/plosone/ .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.