IDEAS home Printed from https://ideas.repec.org/p/arx/papers/2606.06089.html

Leveraging LLMs for Unstructured Claims Data Analysis

Author

Listed:
  • Robert D. Lieberthal

    (Lieberthal and Associates, LLC)

  • Richard Tran

    (MDSight, LLC)

  • Vietbao Phan

    (Thomas Jefferson University)

  • Jawand Singh

    (Lieberthal and Associates, LLC
    William and Mary University)

  • Elizabeth Sottung

    (Thomas Jefferson University)

Abstract

Actuaries rely primarily on structured numerical data for reserving and ratemaking, while valuable predictive information in unstructured text including medical records, adjuster notes, and call transcripts remains largely unused. Manual processing of these documents is time-consuming, inconsistent across reviewers, and unscalable. We present a proof-of-concept framework using large language models (LLMs) to extract structured actuarial variables from unstructured claims data. We implement a two-stage processing architecture separating document-level extraction (Stage 1) from claim-level synthesis (Stage 2). A modular four-script Python pipeline processes synthetic FHIR-based claims data and real claims documents, extracting 36 actuarial variables across reserving, ratemaking, and claims management categories. We validate 14 core variables using two independent clinical expert reviewers scoring 20 synthetic claims on a five-point Likert rubric, achieving mean scores above 4.0 and a weighted kappa of 0.53. Integration with chain ladder reserving demonstrates practical actuarial value: severity-segmented analysis reduced reserve estimation error from 6.5% to 4.0%. The open-source implementation includes audit trails and confidence scoring, providing a replicable foundation for LLM-based actuarial variable extraction in property-casualty insurance.

Suggested Citation

  • Robert D. Lieberthal & Richard Tran & Vietbao Phan & Jawand Singh & Elizabeth Sottung, 2026. "Leveraging LLMs for Unstructured Claims Data Analysis," Papers 2606.06089, arXiv.org.
  • Handle: RePEc:arx:papers:2606.06089
    as

    Download full text from publisher

    File URL: https://arxiv.org/pdf/2606.06089
    File Function: Latest version
    Download Restriction: no
    ---><---

    References listed on IDEAS

    as
    1. Balona, Caesar, 2024. "ActuaryGPT: applications of large language models to insurance and actuarial work," British Actuarial Journal, Cambridge University Press, vol. 29, pages 1-1, January.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Simon Hatzesberger & Iris Nonneman, 2025. "Advanced Applications of Generative AI in Actuarial Science: Case Studies Beyond ChatGPT," Papers 2506.18942, arXiv.org, revised Jun 2026.

    More about this item

    NEP fields

    This paper has been announced in the following NEP Reports:

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:2606.06089. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: https://arxiv.org/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.