IDEAS home Printed from https://ideas.repec.org/a/plo/pone00/0353155.html

Evaluating large language model performance in Risk of Bias assessments: A cross-sectional validation study

Author

Listed:
  • Siddharth Gandhi
  • Arveen Shokravi
  • Yashan Chelliahpillai
  • Michael Balas

Abstract

Objective: To evaluate the reliability and diagnostic performance of ChatGPT-o3 in conducting Risk of Bias (RoB) assessments of randomized clinical trials (RCTs) using the Cochrane RoB 2.0 tool. Materials and methods: This methodological validation study analyzed 50 RCTs sampled from 50 published meta-analyses. Each trial was independently assessed by the original systematic review authors (OSRAs), our masked human panel, and ChatGPT-o3. Structured prompts based on RoB 2.0 guidelines were used to elicit ChatGPT-o3 assessments. Agreement was evaluated using weighted Cohen’s kappa and Gwet’s AC2. Diagnostic performance was measured by sensitivity, specificity, and balanced accuracy, with human ratings as the reference. Results: ChatGPT-o3 classified 34% of trials as high risk, compared with 22% by our panel, and 12% by the OSRAs. Agreement was modest (median κ: 0.33 with our panel; 0.14 with OSRAs). Overall Gwet’s AC2 was 0.30. For detecting high-risk trials, ChatGPT-o3 achieved a sensitivity of 0.46, specificity of 0.69, and balanced accuracy of 0.57. For low-risk trials, its sensitivity was 0.47, specificity was 0.86, and balanced accuracy was 0.66. Discussion: The results indicate that ChatGPT-o3 produced more conservative RoB ratings than human reviewers, identifying a greater percentage of trials as having a high RoB. Conclusion: While unsuitable to be used as a sole assessor, ChatGPT-o3 may serve as an adjunct tool to enhance the efficiency and consistency of RoB assessments in systematic reviews.

Suggested Citation

  • Siddharth Gandhi & Arveen Shokravi & Yashan Chelliahpillai & Michael Balas, 2026. "Evaluating large language model performance in Risk of Bias assessments: A cross-sectional validation study," PLOS ONE, Public Library of Science, vol. 21(7), pages 1-14, July.
  • Handle: RePEc:plo:pone00:0353155
    DOI: 10.1371/journal.pone.0353155
    as

    Download full text from publisher

    File URL: https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0353155
    Download Restriction: no

    File URL: https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0353155&type=printable
    Download Restriction: no

    File URL: https://libkey.io/10.1371/journal.pone.0353155?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pone00:0353155. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: plosone (email available below). General contact details of provider: https://journals.plos.org/plosone/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.