Skip to content
Preprint

Global reanalysis from observations alone with machine learning

Jul 2026 · 0 citations · 43 references
Physics

TL;DR

Results from a prototype system that suggest that observation-trained machine learning models trained only on Earth system observations can be used to generate multi-decade global reanalyses without using physics-based numerical models.

Abstract

Earth system reanalysis datasets are foundational for weather and climate research and provide the gridded training data used by most machine learning weather prediction systems. Here we show results from a prototype system that suggest that machine learning models trained only on Earth system observations can potentially be used to generate multi-decade global reanalyses without using physics-based numerical models. The resulting gridded fields capture large-scale atmospheric structure and variability across multiple timescales, while exhibiting signs of physical coherence in several key dynamical diagnostics. Evaluations of the prototype against held-out independent atmospheric observations indicate that the root mean square vector error of upper-level winds is close to that of ERA5 when compared at a consistent resolution, and that the standard deviation of the error at the surface is between that of 4th- and 5th-generation ECMWF reanalyses (ERA-Interim and ERA5). Furthermore, while traditional reanalysis production is computationally expensive, typically taking several years to produce, the reanalysis presented here was generated during the course of a single working day. These results suggest that observation-trained machine learning models offer a promising new approach for reanalysis production from observations alone.

View source

Similar papers

Preprint Aug 2026

Long-window 4DVar for reanalysis using a differentiable weather model

Atmospheric reanalyses combine observations with model forecasts using complex data assimilation systems. We test whether a differentiable weather model permits a simpler and more accurate method based on a long-window four-dimensional variational data assimilation (4D-Var) formulation that omits the conventional background-error term. The method uses automatic differentiation to find optimal NeuralGCM initial conditions that minimize the misfit to real surface-pressure observations distributed across overlapping windows of two to seven days, assuming no model error. Cycling at 6-hour intervals for three months beginning 1 January 2015 yields a stable reanalysis with smaller error relative to ERA5 in 500-hPa geopotential height than the Twentieth Century Reanalysis version 3 (20CRv3), which uses an ensemble Kalman filter to assimilate the same observations. Every window produces smaller errors than 20CRv3, with analysis error for the four-day window approximately 55% smaller than for 20CRv3. At the end of the four-day window, which does not benefit from future observations, error remains approximately 38% smaller than 20CRv3. Analyses degrade slightly beyond four days, which we attribute to the increasing importance of model error.

Gregory J. Hakim, Jeffrey S. Whitaker, Bo Huang et al. · 0 citations
Jul 2026

Development of Gridded Innovations and Observations Data for Reanalysis Diagnostics

Reanalysis users presently have no clear way to determine which observations have been assimilated at specific points in space and time. Furthermore, for those grid points where observations have been assimilated, users lack efficient quality control metrics to use in their interpretation of the reanalyses. In addition, observation-space data formats such as BUFR have not been deemed the most convenient or efficient to use by reanalysis users because of the complexity and availability of their libraries. To aid reanalysis users and their research, we propose an accompanying gridded data set based on the observations assimilated during the Modern Era Retrospective-analysis for Research and Applications Version 2 (MERRA-2) project. This new dataset is referred to as the MERRA-2 Gridded Innovations and Observations (GIO) and provides the assimilated observations along with the key statistics produced during the observational data assimilation, including (a) the mean forecast departure, (b) the standard deviation of the forecast departure, (c) the number of data counts as well as (d) the bias corrections for satellites radiances. To create GIO data, the observations and innovations are binned to a grid similar to that of MERRA-2 and saved in a convenient NetCDF file. This dataset provides a resource for teaching data assimilation and enables systematic evaluation of observing system impacts in reanalyses, as well as offering training data for machine learning and artificial intelligence applications.

N. Boukachaba, M. Bosilovich, A. D. da Silva et al. · 0 citations
Preprint Aug 2026

Bridging short- and medium-range weather forecasting with machine learning

The National Oceanic and Atmospheric Administration (NOAA) employs independent prediction systems for distinct forecast products. While some separation is practical, we argue that combining short- and medium-range weather into a single prediction system would provide the public with a useful distillation of global weather and its impacts. To this end, we present Nested-EAGLE (Experimental Artificial intelligence Global and Limited-area Ensemble): a 0.25{\deg} global weather model with a 6 km refinement over the Contiguous United States (CONUS). The model achieves significantly lower mean-squared error in near-surface and low-level quantities over CONUS compared to NOAA's Global Forecast System and High-Resolution Rapid Refresh (HRRR), while remaining competitive throughout the rest of the global atmosphere. We show that the skill gains for near-surface fields stem from incorporating high-resolution regional analysis data into training through the nesting process. Forecasts of precipitation amounts are less skillful than those from HRRR, owing to deterministic training. However, we show that Nested-EAGLE provides the most accurate forecasts of storm locations at longer leads, despite blurred extrema. Our results motivate future work to extend the skill gains beyond CONUS and improve precipitation representation.

Timothy A. Smith, Mariah Pope, Sergey Frolov et al. · 0 citations
Preprint Aug 2026

Deep Learning-Based Statistical Downscaling of Sea Surface Temperature Using a Residual Corrective Neural Network

This study proposes a novel deep learning framework that uses a U-Net to generate an initial high-resolution SST estimate, which is subsequently refined using a residual corrective approach, and progressively refines initial U-Net predictions by incorporating dynamically scaled residuals at each step, enabling accurate capture of broad patterns and fine-grained features such as eddies and fronts.

Onkar Jadhav, Tim French, I. Janeković et al. · 1 citation
Sep 2025

Model Training, Data Assimilation, and Forecast Experiments with a Hybrid Atmospheric Model

This paper investigates the performance of a unique proof-of-concept hybrid model in a cycling data assimilation scheme. This previously published model combines the Simplified Parameterization, primitive-Equation Dynamics model (SPEEDY) with an ML-based component that itself is capable of modeling the global atmospheric dynamics. Analysis and forecast experiments are carried out assuming that ERA5 reanalyses, interpolated to the model grid, represent the “true” spatiotemporal evolution of the atmosphere. Six-hourly simulated observations are generated for a 30-year training period and a one-year testing period by randomly perturbing the “true” states. To investigate the effect of the training data on the model performance, the model is trained on different data sets in the different experiments: the training data are either ERA5 reanalyses, analyses prepared using SPEEDY for cycling, or analyses prepared using the hybrid model for cycling. The simulated observations are assimilated with a Local Ensemble Transform Kalman Filter (LETKF) and the length of the ensuing forecasts is 10 days in all experiments. The cycled LETKF remains stable for the entire testing period in all experiments. When the hybrid model is trained on ERA5 reanalyses, the biases of the analyses are negligible and the variance of the analysis error is greatly reduced compared to the experiment in which SPEEDY rather than the hybrid model is used for cycling. The gains in analysis accuracy are more modest when the hybrid model is trained on analyses obtained with SPEEDY or a prior trained version of the model. All forecasts with the hybrid model are more accurate than with SPEEDY.

D. Elliott, Troy Arcomano, I. Szunyogh et al. · 0 citations
Open access Aug 2026

Multi-decadal high-resolution historical simulations with the ECMWF global atmosphere model

This paper describes an ensemble of global atmosphere reference simulations covering the period 1980 to 2023 produced with the European Centre for Medium-Range Weather Forecasts (ECMWF) Integrated Forecasting System (IFS). The resulting dataset consists of 6-hourly three-dimensional global outputs from higher- and lower-resolution configurations of the IFS, which have average horizontal grid spacings of ~9 km (one ensemble member) and ~28 km (ten ensemble members), respectively. The atmosphere is constrained by high-resolution satellite-based estimates of daily mean sea surface temperature and sea ice concentration and time-evolving external climate forcings, including observed greenhouse gas and estimated aerosol concentrations. These data were produced as part of the European Eddy-Rich Earth System Models (EERIE) project and will serve as a reference for corresponding ocean-atmosphere coupled simulations and idealised sensitivity experiments. We have made this dataset publicly available and expect it to be a valuable resource for scientific research and other applications that extend beyond the lifetime of the EERIE project.

M. Aengenheyster, Christopher D. Roberts, Razvan Aguridan et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.