Author
Listed:
- Afrooz Arzehgar
(Mashhad University of Medical Sciences (Iran, Mashhad) - MUMS)
- Saeed Varasteh Yazdi
(EM - EMLyon Business School)
- Hamid Ahanchian
(Mashhad University of Medical Sciences (Iran, Mashhad) - MUMS)
- Saeid Eslami
(UvA - Universiteit van Amsterdam = University of Amsterdam)
Abstract
Early diagnosis of inborn errors of immunity (IEIs) can make a difference in patient outcomes and even cut healthcare costs. However, there are some challenges to overcome, such as clinical complexity, low awareness, and limited resources. Generative artificial intelligence has attracted considerable global attention in medical domains, particularly when integrated into clinical decision support systems (CDSS), as it has the potential to facilitate data interpretation, clinical reasoning, and the optimal use of knowledge resources. Preliminary studies have explored the potential of large language models (LLMs) in various information retrieval tasks, but a systematic evaluation of LLMs with and without retrieval mechanisms for IEI classification is still unexplored. We evaluated and compared the validity and reliability of the responses generated by four open-source and closed-source LLMs, in their baseline form and with augmented data, across 169 IEI patient records, using two input scenarios and four prompt templates. Our primary finding was that the models varied in terms of reliability and performance. The most reliable models were Gemini-1.5-Pro and Llama-3.1-8B-Instruct (K = 0.98) and the best-performing model without data augmentation was Gemini with an F1 score of 43.39 % ± 0.10. The results also showed that retrieval strategies improved the average classification performance, increasing the F1 score from 34 % to 53 % across all models. DeepSeek-R1, which reasoned over retrieved information through the integration of quality refinement and structured retrieval, achieved the best weighted F1 score of 66.94 % 1.19. The study highlights the effective use of generative AI and retrieval-augmented models as a decision support tool for IEI classification. However, incorporating retrieval systems into clinical decision-making processes requires adequate input, effective prompt engineering, and the adoption of retrieval strategies.
Suggested Citation
Afrooz Arzehgar & Saeed Varasteh Yazdi & Hamid Ahanchian & Saeid Eslami, 2026.
"Retrieval-Augmented Language Models for Clinical Decision Support in the Classification of Inborn Errors of Immunity,"
Post-Print
hal-05656789, HAL.
Handle:
RePEc:hal:journl:hal-05656789
DOI: 10.1007/s10875-026-02035-9
Note: View the original document on HAL open archive server: https://hal.science/hal-05656789v1
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:hal:journl:hal-05656789. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: CCSD (email available below). General contact details of provider: https://hal.archives-ouvertes.fr/ .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.