Abstract:To address the limitations of traditional single models when dealing with complex water quality prediction and other real-world challenges, this study proposed the model simple averaging method and two Stacking integration strategies, namely the weighted integration based on the Least Square (LS) method and the Stacking integration with the Extra Trees Regressor (ETR) as the meta-learner. The full-process hourly monitoring data of a certain drinking water treatment plant were collected on-site, and the original data set was subjected to strict preprocessing. On this basis, the predictive performance of all individual models and ensemble models was comprehensively evaluated and compared on the test set. Results indicate that the three proposed ensemble models consistently outperform all single models across all four water quality indicators. Notably, these models achieve a substantial reduction in prediction error. Among these, the LS-weighted model performs the best in terms of average error, with an MAPE value of approximately 1.08%, while the ETR-weighted model has an average error that is slightly higher by 0.02%. Furthermore, this study investigates the weight allocation of base learners in the two Stacking models, revealing that the ETR-weighted model demonstrates greater adaptability. Specifically, this model demonstrates the ability to respond rapidly to complex water quality variations through nonlinearly dynamic weight adjustment, particularly during high-risk events such as sudden turbidity spikes. This study provides a new perspective for the intelligent management of drinking water treatment. The proposed Stacking ensemble model can provide reliable data support for the stable operation and efficiency optimization of water treatment plants. Its strong adaptability offers crucial risk warnings for process disturbances and water quality anomalies, facilitating timely adjustments in subsequent treatment processes.