Author
Listed:
- Malini Deepak
(School of Energy, Geoscience, Infrastructure and Society, Heriot-Watt University, Dubai Campus, Dubai Knowledge Park, Dubai P.O. Box 501745, United Arab Emirates)
- Rabee Rustum
(School of Energy, Geoscience, Infrastructure and Society, Heriot-Watt University, Dubai Campus, Dubai Knowledge Park, Dubai P.O. Box 501745, United Arab Emirates)
Abstract
The activated sludge process is pivotal in wastewater treatment, with ongoing research into its process control methods. Modeling treatment plants aids in analyzing relationships among variables, supporting fault detection and operational decision-making. However, datasets from real-world treatment plants often contain outliers and missing values due to sensor faults, maintenance activities, and operational disruptions, making outlier handling and data imputation essential for reliable modeling. Existing studies on data imputation for activated sludge systems are often based on synthetic or short datasets, limited method comparisons, or inconsistent evaluation metrics, which reduces their applicability to full-scale operational settings. This study addresses these limitations by presenting a comprehensive, head-to-head comparison of Kohonen Self-Organising Maps (KSOM) with widely used multiple imputation and tree-based methods, namely Amelia II, MICE, missForest, and missRanger. The methods are applied to a real-world multivariate dataset comprising 19 process variables collected over 8.5 years from a full-scale activated sludge treatment plant, containing 39% overall missing data with highly uneven missingness across variables. A validation framework based on held-out observation data is used, and performance is assessed using complementary metrics, including the coefficient of determination (R 2 ), average absolute error (AAE), relative average absolute error (RAAE), mean squared error (MSE), and root mean squared error (RMSE). Results show that KSOM consistently outperforms the competing methods across most variables and evaluation metrics. KSOM achieves near-perfect R 2 values (≈1) for many process variables, with lower absolute and relative errors, even for variables with very high (>70%) and irregular missingness. These findings highlight KSOM’s robustness in capturing multivariate relationships and cluster structure in complex, operational WWTP data.
Suggested Citation
Malini Deepak & Rabee Rustum, 2026.
"Comparison of Imputation Methods for Activated Sludge Data: A Case Study on Imputing Missing Data,"
Waste, MDPI, vol. 4(2), pages 1-24, May.
Handle:
RePEc:gam:jwaste:v:4:y:2026:i:2:p:17-:d:1953675
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:gam:jwaste:v:4:y:2026:i:2:p:17-:d:1953675. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: MDPI Indexing Manager The email address of this maintainer does not seem to be valid anymore. Please ask MDPI Indexing Manager to update the entry or send us the correct address
(email available below). General contact details of provider: https://www.mdpi.com .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.