Shafi, Sedra (2025) Machine Learning-Based Reconstruction of Air Pollution Data in South-East and Western Pacific Asia. [Tesi di dottorato]

[thumbnail of Thesis_sedrashafi_38cycle.pdf] Documento PDF
Thesis_sedrashafi_38cycle.pdf

Download (12MB)
Tipologia del documento: Tesi di dottorato
Lingua: English
Titolo: Machine Learning-Based Reconstruction of Air Pollution Data in South-East and Western Pacific Asia
Autori:
Autore
Email
Shafi, Sedra
sedra.shafi@unina.it
Data: 10 Dicembre 2025
Numero di pagine: 157
Istituzione: Università degli Studi di Napoli Federico II
Dipartimento: Scienze della Terra, dell'Ambiente e delle Risorse
Dottorato: Scienze della Terra
Ciclo di dottorato: 38
Coordinatore del Corso di dottorato:
nome
email
Ferranti, Luigi
lferrant@unina.it
Tutor:
nome
email
Scafetta, Nicola
[non definito]
Data: 10 Dicembre 2025
Numero di pagine: 157
Parole chiave: Keywords: Air pollution (PM2.5, PM10, O3, NO2, SO2, and CO) modeling; Machine learning regression models; Air pollution (PM2.5) assessment; Meteorological monsoon conditions; South-East and Western-PacificAsia; Reconstruction of missing data; The city of Delhi, India,Temperature, Missing data of Naples, 10 Stations
Settori scientifico-disciplinari del MIUR: Area 04 - Scienze della terra > GEO/12 - Oceanografia e fisica dell'atmosfera
Informazioni aggiuntive: 38 cycle
Depositato il: 23 Dic 2025 08:45
Ultima modifica: 12 Ago 2026 05:37
URI: https://www.fedoa.unina.it/id/eprint/16115

Abstract

The rapid deterioration of air quality across South-East and Western Pacific Asia is accelerating due to population growth and industrial expansion. Anthropogenic activities — particularly rapid urbanization — have significantly increased pollution levels. Air pollutants such as PM2.5, PM10, O₃, NO₂, SO₂, and CO pose serious environmental and public health challenges, especially in densely populated metropolitan areas like Delhi. The substantial rise in fine particulate matter (PM) represents a critical threat to human health, contributing to respiratory and cardiovascular diseases and further degrading air quality. Exposure levels in fast-developing regions are relatively high due to unsynchronized urbanization processes. Although these areas suffer from severe air pollution, detailed and comprehensive data remain scarce. Meteorological factors — such as wind speed, barometric pressure, temperature, rainfall, and monsoon seasonality — play a significant role in influencing pollution levels, particularly PM concentrations. However, accurate hazard assessment requires long-term, complete, and spatially distributed monitoring data, which are often lacking. For example, the World Air Quality Historical Database listed only four operational monitoring stations in Delhi in 2014; this number gradually increased to 45 by 2024. Such disparities result in inconsistent and incomplete records, hindering efforts to evaluate urban air quality, assess policy effectiveness, and manage pollution levels. Research Objectives. The first objective of our research is to apply a statistical modeling approach to estimate daily PM levels based on meteorological parameters in five major polluted cities: Lahore (Pakistan), Delhi (India), Dhaka (Bangladesh), Hanoi (Vietnam), and Shanghai (China). The second objective is to reconstruct missing daily air pollution data — specifically PM2.5, PM10, O₃, NO₂, SO₂, and CO — from 2014 to 2024 across 45 monitoring stations in Delhi. We propose a machine learning (ML)-based statistical reconstruction of six daily atmospheric pollution indicators to build a more consistent and spatially representative database for the entire city over the 11-year period. This network is then used to generate ensemble average records that more accurately reflect the daily evolution of air pollution concentrations in Delhi since 2014. Third objective is the reconstruction of missing temperature data in Naples, across 11 monitoring stations from 2011 to 2025 by using Monte Carlo–based machine learning framework. Also we reconstructed the non-operational period (2022–2025) of the S. Marcellino_1 station and then compared the reconstructed data with the actual observations from the operational S. Marcellino_2 station to assess the model’s accuracy and reliability. The goal was to identify an optimal model that balances precision with computational cost for large-scale meteorological datasets. Methodology. Our research employs a Monte Carlo-driven machine learning (ML) modeling framework, involving a comparative analysis of 35 different ML regression techniques. The goal is to identify the most effective algorithms for reconstructing and predicting PM levels using meteorological variables alone, and to statistically cross-reconstruct missing daily pollution records. In the first study, each ML regression model was trained using data from 2020–2021 to reconstruct daily PM levels. These models were then used to fill in missing data at three-year intervals and to forecast PM levels for the entire year of 2022 using only meteorological inputs from that year. The second study introduces a methodology for statistically cross-reconstructing missing daily pollution records for all six pollutants across all 45 stations from 2014 to 2024. For each station with missing data, the algorithm identifies its four most correlated pollution records from other regions of the city—assuming similar environmental conditions—and uses these as inputs in MATLAB’s Regression Learner tool to test 35 ML regression techniques. The best-performing models are selected based on their reconstruction accuracy and used to fill all possible data gaps. This iterative process continues until all missing records are addressed. In third study we applied the same Monte Carlo–based machine learning framework discussed in the second study to the temperature in climatic zonation of Naples, Italy. We assesedmultiple regression models, including Linear Regression, Stepwise Linear Regression, and other Machine Learning models like Fine Tree (Regression Trees) and Bagged Tree (Ensembles of Trees) to assess their performance against the observed data. Results and Insights. The results of the first study indicate that most daily and seasonal variability in PM levels can be effectively reconstructed from meteorological data. However, performance varied significantly across ML models, as measured by Root Mean Square Error (RMSE) tests. The Ensemble Boosted Tree method demonstrated optimal efficiency during the training period (2020–2021) and was highly effective in predicting PM levels for 2022. Additionally, the Trilayer Neural Network model proved most suitable for reconstructing short-term missing data after three years of training. In contrast, traditional multi-linear regression models consistently underperformed in both reconstruction and prediction tasks. Our comparative analysis underscores the importance of evaluating multiple ML regression methodologies to identify those best suited for reconstructing PM records from meteorological parameters. For the second objective, the most effective ML algorithms included Fine Tree, Bagged Trees, Optimizable Ensemble, Fine Gaussian SVM, Rational Quadratic, and Exponential models. These outperformed traditional regression approaches due to their ability to capture complex non-linear relationships among variables. The reconstructed datasets enabled the generation of ensemble mean records for each pollutant, offering a more realistic representation of Delhi’s daily pollution dynamics and correcting biases in long-term trends caused by non-homogeneous data. For the third objective, The performance of Linear Regression model the Stepwise Linear Regression model demonstrated the best performance, achieving correlation coefficients of 0.9989 (yearly) and 0.9962 (daily) between the reconstructed data of S. Marcellino 1 and the measurements from S. Marcellino 2, slightly outperforming the standard Linear Regression (0.9988 and 0.9961). However, its computational cost was almost ten times higher (1391 s vs 146 s). The other methodologies yielded slightly lower correlations by requiring significantly longer processing times. However, Stepwise Linear Regression correlation provided the most reliable, fast and time efficient results for this dataset. Broader Implications. This dissertation highlights the innovative application of ML regression techniques for forecasting air quality and reconstructing missing pollution data — an essential tool for policymaking in South-East and Western Pacific Asia, where pollution monitoring infrastructure remains limited. The final ensemble results reveal modest improvements in air quality between 2014 and 2024. Beyond Delhi and the other cities here analyzed, this statistical approach can be effectively applied to other urban environments and scenarios — as shown in Chapter 5 where the case of Naples is briefly analyzed — offering a powerful framework for studying air conditions in complex cities with incomplete or fragmented data. Also methodology offers a practical and reliable framework for enhancing the continuity of climate datasets and can be readily applied to other meteorological variables, including pressure, dew point, humidity, and wind speed.

Downloads

Downloads per month over past year

Actions (login required)

Modifica documento Modifica documento