Sensor Fusion and Machine Learning for Real-world Emissions Monitoring and Climate-related Environmental Assessment in the United States: A Critical Review
Aug 2026· Asian journal of current research· Vol 11, pp. 479-504· 0 citations
TL;DR
Priorities identified include standardised blind evaluation for ambient sensing, detection-probability-aware fusion, distribution-shift-resilient calibration, and reporting practices that expose spatial error structure rather than aggregate it away.
Abstract
Emissions monitoring in the United States is being reconstructed around dense, heterogeneous sensing and algorithmic inference. Low-cost electrochemical and optical sensors, mobile platforms, aircraft-mounted imaging spectrometers, point-in-space continuous monitors, and polar-orbiting and geostationary satellites now generate observations at scales and frequencies that regulatory reference networks were never designed to provide, and machine learning increasingly mediates the step from raw signal to reported emission or exposure. This review critically appraises what that reconstruction has demonstrated, what remains uncertain, and where confidence is currently misplaced. Evidence was assembled through registry-based bibliographic searching and verified source-by-source, with emphasis on studies conducted in United States settings and published from 1995 onwards. Four findings recur across otherwise separate literatures. First, reported model performance is dominated by co-location conditions rather than deployment conditions, and cross-site, cross-season and cross-instrument transfer degrades in ways that headline coefficients of determination conceal. Second, blind controlled-release testing has become the strongest evidentiary design available for methane sensing and has no established equivalent in ambient sensor calibration, producing a marked asymmetry in the quality of validation evidence between adjacent fields. Third, heavy-tailed and intermittent emission distributions break the sampling assumptions embedded in both survey design and supervised learning, so that platform-specific estimates diverge for structural rather than analytical reasons. Fourth, fusion across platforms with different detection thresholds is frequently treated as additive when it is in fact a partial-detection problem requiring explicit statistical treatment. Errors in fused products are spatially structured and correlate with monitor density, so that inferential accuracy is lowest where policy attention is greatest. Priorities identified include standardised blind evaluation for ambient sensing, detection-probability-aware fusion, distribution-shift-resilient calibration, and reporting practices that expose spatial error structure rather than aggregate it away.
The convergence of next-generation sensors, sophisticated data analytics, and Earth system models has fundamentally transformed the landscape of environmental radiation monitoring. These technological synergies have collectively enhanced the accuracy, spatial coverage, and real-time responsiveness of detection systems, enabling more precise identification and assessment of radiation exposure risks. This article systematically reviews recent progress in sensor technologies, including improved sensitivity and durability, alongside the growing role of machine learning algorithms in processing complex monitoring data and distinguishing anomalous radiation patterns from background noise. A central focus is the integration of radiation data into Earth system models, which facilitates the simulation of radionuclide dispersion and environmental fate, thereby supporting predictive capabilities and informed decision-making during radiological emergencies. Furthermore, the article highlights the practical benefits of real-time monitoring networks and early warning systems in strengthening disaster preparedness and safeguarding public health. Despite these notable advances, significant challenges remain, including issues related to sensor calibration, cross-platform data integration, infrastructure deployment, and cybersecurity. Looking ahead, continued research, technological innovation, and international collaboration are essential to enhance the reliability, scalability, and global interoperability of these systems. Ultimately, these developments not only contribute to reducing radiation-related risks but also align with broader sustainability objectives, particularly in supporting environmental health, climate resilience, and evidence-based policy formulation under the United Nations Sustainable Development Goals.
Wenbing Wang· Journal of Environmental &am...· 0 citations
Emerging as the new frontier of environmental monitoring and environmental science studies, artificial intelligence (AI) promises to allow the process of scalable inferences when new volumes of satellite, airborne, in situ, and model-generated data are taken into account. This review brings together the data ecosystems, methodological principles and application evidence that shape the new frontier of Earth observation, namely the digital frontier. Measurement physics: We explain the effects of measurement physics on problem formulation, label uncertainty, and missingness, and how current machine-learning practices are naively transferred to other domains, despite these domains exhibiting different possibilities that could affect model performance. After this, we discuss principal AI strategies focusing on representation learning and self-supervised pretraining, spatio-temporal deep learning in map and prediction, multi-modal fusion, and generative learning in gap filling, downsizing, and reconstruction. Specific focus is made on physics-guided and hybrid modeling approaches that jointly integrate learned components with mechanistic models to enhance plausibility, extrapolation and uncertainty quantification, calibration, and interpretability needed to gain scientific credibility and operational decision support. In the fields of application that we have considered, land systems, atmosphere and air quality, hydrology and water resources, cryosphere, ocean and coasts, natural hazards and urban environments, we discuss common patterns of success and failure, with operational readiness spanning almost equally evaluation design, data governance, lifecycle maintenance, and architecture choice. Our final contribution is research and community priorities such as Earth system foundation models, resilient extremes and out of distribution beneficial products, decision facing probabilistic products and responsible governance that deal with bias, privacy and dual-use risks. This combination of directions defines a roadmap on the way to credible prototypes to reliable and reproducible and beneficial Earth AI systems.
Xuebin Wang, Huanle Zhang· Journal of Environmental &am...· 0 citations
Accurately predicting fine particulate matter (PM2.5) concentrations in regions with sparse monitoring networks remains a critical challenge for air quality management and public health. This study evaluates a machine learning (ML) data fusion approach that integrates daily federal regulatory observations, daily low-cost community sensor measurements, and monthly satellite-derived aerosol products (functioning as a regional background field) to improve PM2.5 prediction across under-monitored environments. Using a Long Short-Term Memory (LSTM) neural network architecture, the analysis examines how combining heterogeneous data sources influences predictions. Results show that pooled multi-source training was associated with higher holdout skill relative to some single-source configurations under this parsimonious baseline, though associations are city- and configuration-dependent and cannot be attributed solely to fusion because evaluation populations are not common. Comparisons against tree-based baselines (Random Forest, Gradient Boosting, XGBoost) indicate that overall predictive skill, not just the LSTM’s, is constrained by data availability, suggesting that data composition, rather than model choice, is the primary driver of the observed performance patterns. These findings highlight both the potential and the practical constraints of multi-source ML approaches for air quality prediction and exposure assessment, with implications for model design, monitoring strategy, and environmental equity. This study is intentionally scoped as an applied evaluation of data fusion performance rather than a comprehensive assessment of algorithmic optimality or operational forecasting readiness. The analysis focuses on daily PM2.5 prediction across a selected set of U.S. cities and does not address sub-daily variability, real-time deployment constraints, or event-specific model optimization. Model performance is therefore interpreted in the context of data availability, consistency, and representativeness, rather than as an upper bound on achievable predictive skill.
Suhrudh Chivukula, Adrian J. Cortes Santos, R. Delgado et al.· Atmosphere· 0 citations
Studying climate change requires reducing uncertainties in CO2 and CH4 emission estimates to better distinguish anthropogenic from natural sources, which motivates spaceborne measurements with improved revisit frequency and spatial coverage. In this context, the Horizon Europe SCARBOn project assesses a low-cost satellite constellation featuring the NanoCarb imaging interferometer as its core sensor for monitoring CO2 and CH4 emissions in the atmosphere. However, estimating CO2 and CH4 concentrations with high revisit and spatial coverage poses significant challenges: full-physics retrieval algorithms commonly used rely on repeated high-resolution radiative transfer (RT) simulations, which are computationally expensive when using line-by-line RT models. As an alternative, we propose in this study a feedforward multilayer perceptron (MLP) surrogate designed to accurately and efficiently predict top-of-atmosphere radiances in the CO2 weak band, using a combined mean absolute error (MAE) loss on radiances and RT Jacobians to preserve both spectral accuracy and sensitivity to geophysical parameters. Coupling the MLP-based RT surrogate with the NanoCarb instrumental response yields an efficient and precise forward model for NanoCarb measurements, which shows promising results for CO2 concentration retrieval.
Jordan Lontsi Tedongmo, Y. Ferrec, Laurence Croizé et al.· 0 citations
Recent years have seen a rapid expansion in the production of large-scale geospatial maps derived from Earth observation (EO) data, driven largely by advances in machine learning (ML) and large computing infrastructure. Although the barrier to generating such maps has dropped substantially, established best practices have yet to emerge, and design decisions made early in the pipeline can quietly propagate errors into the final product. Producing a technically sound and scientifically credible product remains challenging. Choices made at every stage are tightly coupled: preprocessing decisions shape the training signal, dataset design governs what the model can learn and how reliably its performance can be assessed, and global-scale inference introduces engineering challenges in compute and data access at scale, as well as artifact mitigation. Furthermore, uncertainty quantification and independent map validation each require dedicated methodological attention that is often underestimated. This paper presents a concise, end-to-end account of the recommended practices spanning the pipeline from satellite data to an operational map product. We organize the discussion around six interconnected themes: the EO data infrastructure landscape, data selection and preprocessing, ML dataset construction and model training, uncertainty quantification, map production and distribution, and validation. This paper is a condensed version of a longer guide that provides greater depth across all stages, accessible online at ghjuliasialelli.github.io/ML-EO-Maps/.
Ghjulia Sialelli, Robin Young, Yu-Chang Jiang et al.· arXiv.org· 1 citation
Dissolved oxygen (DO) is the single most informative indicator of the ecological state of running waters, yet the optical probes used to record it are also the most prone to fouling, drift and outright failure among the routine sensors deployed at gauging stations. The resulting gaps interrupt exactly the early-warning function that continuous monitoring is meant to serve. This study develops a DO soft-sensor that reconstructs the concentration from the more robust and inexpensive variables recorded alongside it — water temperature and discharge — without recourse to the DO record itself, so that it remains usable while the oxygen probe is offline. A compact set of physically motivated predictors, including the temperature-dependent oxygen-solubility term of Benson and Krause, seasonal harmonics and short-memory hydrological statistics, is passed to a non-negativity-constrained stacked ensemble that combines five tree-based learners through a super-learner meta-model. The framework is evaluated on a nine-year daily record (n = 3357; 2012–2021) from a large lowland river under a strictly chronological train–validation–test partition. On the held-out final 504 days the ensemble attains R² = 0.877, RMSE = 0.697 mg L⁻¹, Nash–Sutcliffe efficiency 0.877 and Kling–Gupta efficiency 0.899, improving on a temperature-only physical baseline (RMSE 0.942 mg L⁻¹) and on multiple linear regression, and performing on par with the best individual learner while providing a single robust predictor. SHAP attribution recovers the correct physical hierarchy — the smoothed water-temperature signal and oxygen solubility dominate, with discharge and seasonality secondary — confirming that the model is physically consistent rather than an opaque fit. Split-conformal prediction intervals deliver near-nominal empirical coverage (0.82, 0.90 and 0.93 against nominal 0.80, 0.90 and 0.95), and the reconstructed series flags low-DO days with a recall of 0.97 and an F₁-score of 0.85. A rolling-origin backtest over four successive out-of-period blocks confirms that this performance is not an artefact of a single split, with a mean out-of-period R² of 0.854. The complete computational pipeline, the exact data-retrieval script, the preprocessed feature matrix, all model predictions with intervals, the SHAP values and the pinned software environment are released so that every table and figure below can be regenerated with a single command.
Lakkireddy Venkateswara Reddy, C. Manjunath, M. Rao et al.· Ecological Engineering &...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.