Skip to content

Software supporting "Spatiotemporal Mapping of Sea Surface Height and Temperature Fields from Satellite Observations Using Score-based Data Assimilation"

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

# Score-based data assimilation of SWOT SSHA and ACSPO SST — code, model weights and derived data > **Note:** this archive is the reference code for the paper's Zenodo record. It is a > snapshot of the research code as it was used to produce the paper's results. It is not a > polished package and should not be treated as one for wider use. We are developing a > usable implementation at . This zenodo record is a companion archive to a submitted manuscript, "Spatiotemporal Mapping of Sea Surface Height and Temperature Fields from Satellite Observations Using Score-based Data Assimilation". It holds the analysis code and notebooks, the trained SSHA and SST diffusion models (with the checkpoint each experiment used), the code and configs that trained them, and the small derived datasets behind the paper's tables and figures. ## Contents | Path | What it is | |---|---| | `model_weights/` | Four checkpoints (NVIDIA Modulus `.mdlus`), two from each of the two trained score networks, and the Hydra config each training run wrote out. See below for which experiment used which checkpoint. | | `training/` | Training code (`train_diff.py`, `training_diff/`), the training data loader (`data_loader/`), the configs for the two runs (`conf/013_*`, `conf/014_*`, plus the shared `config_train_base.yaml`) and their training logs. | | `SDA_scripts/` | Inference, evaluation and figure notebooks, the shared helpers in `src/`, data-fetch scripts in `scripts/`, observation masks, and the derived data listed below. | | `oa_for_tatsu/` | `swot_oa`, the Python objective-analysis (OA) baseline used in the OA appendix, with its tests and diagnostics notebooks. | | `third_party/sda/` | The `sda` package (F. Rozet, MIT licence), which provides the VP SDE and Gaussian score used for guided sampling. | | `environment.yml` | The conda environment. | | `MD5SUMS.txt` | Checksums for every file in this archive. | ### Model weights Two score networks were trained, one per field, each on 5 × 12-hourly steps of llc4320 North Pacific output: `013_SST_5tstep_12hrly` (SST, ran to step 200000) and `014_SSH_5tstep_12hrly` (SSHA, ran to step 154547). Two checkpoints of each run were used: | File | Field | Checkpoint | Used for | |---|---|---|---| | `SST_013_5tstep_12hrly_step117522.mdlus` | SST | 117522 | tile / season sweep | | `SSH_014_5tstep_12hrly_step154495.mdlus` | SSHA | 154495 | tile / season sweep | | `SST_013_5tstep_12hrly_step199995.mdlus` | SST | 199995 | noise-level and observation-coverage tests, real ACSPO observations | | `SSH_014_5tstep_12hrly_step154495.mdlus` | SSHA | 154495 | observation-coverage tests, real SWOT observations | The two tile/season-sweep checkpoints are also in `SDA_scripts/diff_models/`, where the sweep scripts in `SDA_scripts/src/run_*_tile_sweep.py` expect them. The experiment notebooks (`004_*`, `005_*`, `007_*`) and `gen_SST_real_obs_MUR.py` instead pick the second-to-last checkpoint in the training output directory, which after training finished resolves to steps 199995 / 154495. The real-observation SST outputs record the checkpoint they used in their `model` attribute. Normalisation constants used throughout: `std_ssh = 0.04537` m, `mean_ssh = 0`, `std_sst = 5.5` °C (SST is normalised against a seasonal climatology). ### Derived data in `SDA_scripts/` | Path | Used for | |---|---| | `processed_tiles/ensmean_cache/` | Ensemble-mean RMSE score / effective-resolution metrics for the tile–season sweeps and the noise / mask ablations (`010_tilewise_skill_panels.ipynb`). | | `processed_tiles/ensemble_stats_cache/` | Rank-histogram and spread–skill statistics (ensemble-calibration appendix, `009_*`). | | `processed_tiles/OA_baseline/` | OA and SDA skill and spectra cubes for the OA comparison. | | `processed_tiles/SSH_real_obs/`, `processed_tiles/SST_real_obs_MUR/` | Processed real-observation reconstructions (SWOT KaRIn SSHA, ACSPO PM SST). | | `Multi_step_predictions/SSH_real_obs/`, `Multi_step_predictions/SST_real_obs_MUR/` | The raw 30-member real-observation ensembles. | | `real_observations/*_tiles_for_inference*`, `real_obs_tile_geolocation.csv` | The regridded real-observation input tiles and their coordinates. | | `masks/`, `testing_masks/` | SWOT / nadir sampling masks and ACSPO cloud masks for the synthetic experiments. | **Not included** (tens to hundreds of GB): the llc4320 training and testing fields, the full synthetic-sweep ensembles (`processed_tiles/SSH`, `processed_tiles/SST`, the noise- and mask-test cubes), intermediate training checkpoints, and the raw regridded ACSPO / SWOT swaths. Notebooks that read those files are kept with their rendered outputs, so every figure and table can be inspected without re-running. Re-running them needs the full sweep outputs, which are not archived here. ## Experimental settings | Setting | SSHA | SST | |---|---|---| | Time steps per sequence | 8 (12-hourly) | 15 (12-hourly); 8 for real observations | | Ensemble size | 30 | 30 (28 in the tile/season sweep) | | Synthetic observation noise σobs (tile/season sweep) | 0.025 normalised ≈ 0.11 cm | 0.01 normalised ≈ 0.055 °C | | Assumed σobs for real observations | 0.01 m | 0.2 °C (0.036 normalised) | | Guidance γ / diffusion steps / Langevin corrections | 0.01 / 256 / 2 (τ = 0.1) | same | RMSE score and WPSD score are computed over all grid points, both observed and unobserved. For SST the spatial mean is removed from both fields at each time step before scoring. ## Dependencies - Python environment: `environment.yml`. - `sda`: included under `third_party/sda/`. Put it on `PYTHONPATH` or `pip install -e` its parent. - NVIDIA Modulus 0.7.0a0 is needed to load `.mdlus` checkpoints and to run the training code. It is not redistributed here. - Training data loader: `training/data_loader/` holds `claude_data_loaders.py` from the authors' `SWOT-inpainting-DL` repository (commit `34087e2`, 2025-09-25) and the `interp_utils.py` it imports from `SWOT-data-analysis` (commit `54e3b78`, 2025-08-12). These are the last committed versions before training. Later local edits (a path change and an unused cloud-mask option) are not included. The training scripts and the loader add hard-coded absolute paths to `sys.path`. To run them elsewhere, point `sys.path` at `training/data_loader/` instead. ## Not included for licensing reasons The original MATLAB objective-analysis example that `swot_oa` was ported from is not redistributed. `oa_for_tatsu/README.md` documents the port and every deliberate divergence from the MATLAB version.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#computer vision Review Mar 2008

Agile methods in European embedded software development organisations: a survey on the actual use and usefulness of Extreme Programming and Scrum

The results show that the embedded industry has been able to apply agile methods in its development processes and that the appreciation of the agile methods and their individual practices appears to increase once adopted and applied in practice.

O. Salo, P. Abrahamsson · 238 citations · ⚡9
#computer vision Open access Jul 2017

What happens when software developers are (un)happy

Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 236 citations · ⚡13
#computer vision Open access Oct 2004

Mobile-D: an agile approach for mobile application development

The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.

P. Abrahamsson, Antti Hanhineva, H. Hulkko et al. · 225 citations · ⚡18

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.