IDEAS home Printed from https://ideas.repec.org/p/osf/lawarc/9wq32_v2.html

Validating an AI Grader for Community Corrections Supervision Visits: A Reliability and Agreement Study Against Expert Human Auditors

Author

Listed:
  • Johnston, Cliff Hurt
  • Green, Ted
  • Meade, Valerie
  • Warren, Madeline

Abstract

Meta-analytic evidence links sustained officer fidelity to core correctional practices with substantially lower participant recidivism; caseloads supervised by untrained officers recidivate roughly 39% more often than those supervised by trained officers (Chadwick et al., 2015). That fidelity, however, can be measured today only by having trained humans hand-code recordings, which does not scale beyond a small audit sample. We test whether a frontier large language model can grade recorded supervision visits against a rubric of observable, standards-based officer behaviors as reliably as expert human auditors. On a 30-visit corpus drawn from multiple agencies and independently scored by three expert graders, we report a descriptive validation of an LLM grader (Claude Sonnet 5 running a calibrated prompt, scored as a three-run ensemble). The expert graders agree only moderately with one another (item-level agreement 65–73%, Gwet’s AC1 +0.34 to +0.46; weighted-total Krippendorff’s α +0.199), a level typical of subjective behavioral coding. Against this panel, the model lands at least as close to the panel average as the median human grader on 23 of 30 visits (77%, 95% CI 0.60–0.90; single-run range 20–24 of 30; 16 of 18 on the never-seen holdout). It agrees with the panel at the item level about as well as the graders agree with each other (macro-F1 0.57–0.74), sits closer to each grader than the graders sit to one another, and, consistent with sitting near the panel’s center, raises the panel’s agreement statistic when added as a member. Grader-prompt performance is tied to the specific model version. We scope the tool to officer coaching, not personnel decisions or participant outcomes, and discuss deployment, reliability, cost, and limitations, along with next steps: an adjudication study to establish a consensus reference, and a study of whether the model can coach as well as a human.

Suggested Citation

  • Johnston, Cliff Hurt & Green, Ted & Meade, Valerie & Warren, Madeline, 2026. "Validating an AI Grader for Community Corrections Supervision Visits: A Reliability and Agreement Study Against Expert Human Auditors," LawArchive 9wq32_v2, Center for Open Science.
  • Handle: RePEc:osf:lawarc:9wq32_v2
    DOI: 10.31228/osf.io/9wq32_v2
    as

    Download full text from publisher

    File URL: https://osf.io/download/6aa023e599290ba88548627e/
    Download Restriction: no

    File URL: https://libkey.io/10.31228/osf.io/9wq32_v2?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:osf:lawarc:9wq32_v2. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: OSF (email available below). General contact details of provider: https://lawarchive.info/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.