Author
Listed:
- Jiratchaya Nakbang
- Chonthicha Arbsuwan
- Santitham Prom-on
- Peerut Chienwichai
Abstract
Antimicrobial resistance (AMR) is a critical global health problem that has become increasingly alarming in recent years. The discovery of new antibiotics is one approach for alleviating AMR, however, screening for novel drugs is time consuming and expensive. To accelerate antibiotic discovery, the integration of machine learning algorithms with Quantitative Structure–Activity Relationship (QSAR) calculations could provide a rapid solution. Thus, this study combines a QSAR model and machine learning algorithms to predict antibacterial activities of potential novel drugs based on chemical information. Information on compounds that are reportedly active and inactive against bacteria was downloaded from the PubChem database and manually curated to create positive and negative datasets. The decision tree (DT), support vector machine (SVM), and naïve Bayesian (NB) algorithms were employed to predict the antibacterial activities of chemical compounds from their Simplified Molecular Input Line Entry System (SMILES) information. The models were then evaluated quantitatively and tuned. DT and SVM exhibited comparable predictive performance and outperformed the NB model, achieving accuracy, precision, sensitivity, and AUC-ROC values exceeding 0.90. DT was chosen for further analysis because of its simplicity and effectiveness. This revealed that descriptors relating to the electrotopology and β-lactam structures of compounds were the top contributors to model predictability. The model was then further tested against different classes of antibiotics and achieved high accuracy in all classes. The model is freely available as a web application at: https://antibacterial-predictor-model-ocogzqyibervrqb7trvfev.streamlit.app/.Author summary: Antimicrobial resistance (AMR) is a serious global health challenge, with its incidence steadily increasing while the rate of new antibiotic development continues to decline. The process of antibiotic discovery is costly and time-consuming, making it difficult to keep pace with the growing threat of resistant pathogens. To address this urgent need, computer-aided drug discovery has emerged as a promising solution. In this study, we developed a machine learning model for antibiotic prediction. Our model demonstrated strong performance, achieving over 90% accuracy, precision, and sensitivity, effectively distinguishing compounds with antibacterial properties from those without. Importantly, the model’s decision-making process is interpretable, enhancing the reliability of its predictions. Furthermore, we have packaged the model into a web application, making it accessible for broader use. This tool aims to accelerate antibiotic discovery, particularly for researchers without a background in data science, and contributes to the global effort to combat AMR.
Suggested Citation
Jiratchaya Nakbang & Chonthicha Arbsuwan & Santitham Prom-on & Peerut Chienwichai, 2026.
"Machine learning-based quantitative structure–activity relationship model for antibiotic prediction and discovery,"
PLOS Digital Health, Public Library of Science, vol. 5(8), pages 1-16, August.
Handle:
RePEc:plo:pdig00:0001593
DOI: 10.1371/journal.pdig.0001593
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:plo:pdig00:0001593. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: digitalhealth (email available below). General contact details of provider: https://journals.plos.org/digitalhealth .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.