Skip to content
Open access

Continual learning for autoregressive PDE surrogates under evolving physical regimes

Aug 2026 · Machine Learning for Computational Science and Engineering · Vol 2 · 0 citations · 73 references

Abstract

Deploying machine learning surrogates in scientific simulations faces multifaceted challenges, primary among which is the lack of Continual Learning (CL) capabilities—specifically, the inability to adapt to new physical regimes without significantly degrading performance on prior ones. This is particularly problematic for autoregressive surrogates of time-dependent Partial Differential Equations (PDEs), where small prediction errors can accumulate over long rollouts and new physical regimes overwrite previously learned dynamics. We formulate this adaptation as a CL problem, demonstrating that while standard Experience Replay (ER) is a robust baseline across Advection-Diffusion, Burgers’, and Navier-Stokes equations, storing full high-resolution rollouts can be memory-inefficient. To address this, we introduce Replay-TS, a temporal-slicing replay strategy that stores compact autoregressive windows sampled across past simulations. Through empirical analysis, we show that Replay-TS exploits the low-frequency spectral redundancy of physical systems to enable sparse supervision for rollout steps. By preserving the contiguous historical context and sparsely penalizing the autoregressive target steps, Replay-TS improves retention performance under a fixed memory budget by leveraging higher sample diversity. Replay-TS consistently outperforms standard ER methods across standard 1D and 2D streams, achieving over a 30% MSE reduction in a mixed-physics stream, while remaining architecture-agnostic.

Read PDF

Similar papers

Preprint Aug 2026

Distillation of Foundation Models for Time-dependent PDEs

This work proposes Teacher Rollout Extension (TREX), a knowledge distillation framework that transfers the predictive capability of a pretrained foundation model into a compact and efficient student that can match or surpass the teacher's accuracy while reducing the number of parameters by several orders of magnitude and achieving more than an order-of-magnitude speedup in inference.

Daniel Musekamp, Boshra Ariguib, Andrei Manolache et al. · 0 citations
Jul 2026

HERO: History-Enriched Rollout Training for Long-Horizon Autoregressive Neural Operators

Experiments show that HERO consistently improves long-horizon accuracy, stable rollout length, and out-of-distribution robustness at no inference-time cost, indicating that history-enriched relative supervision is effective for stabilizing long-horizon autoregressive prediction.

Jiaquan Zhang, Shuxu Chen, Haifan Meng et al. · 0 citations
Jun 2026

Attention-based Experience Replay Framework for Continual Learning of Agnostic Time Series Forecasting Models

This work introduces a novel framework for continual time series forecasting, designed to extend existing static forecasting models commonly used in the literature by incorporating an Experience Replay strategy guided by Attention mechanisms, which allows the model to adapt dynamically to new contexts while preserving prior knowledge, effectively mitigating catastrophic forgetting.

Quentin Besnard, Nicolas Ragot · 0 citations
Preprint Aug 2026

DAW: Dynamics-Aware Weighting for Deep Learning Forecasts of Chaotic Systems

DAW reshapes the loss landscape to allocate representational capacity toward the sparse, high-d regimes where forecast errors are systematically large, and consistently outperforms uniform training, purely statistical density weighting, and its randomly permuted ablation on the chaotic KS equation.

Fang Zhou, Gianmarco Mengaldo · 0 citations
Preprint Aug 2026

Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulations

This paper proposes Kastor, a comprehensive methodology to adapt a deterministic physics foundation model into a highly efficient and accurate generative surrogate, and introduces a two-stage inference scheme that combines a large-stride causal auto-regressive model with a non-causal temporal super-resolution network, significantly reducing error accumulation while minimizing computational cost.

Guillaume Couairon, Alexis Jacq, Yu-Han Wu et al. · 0 citations
Preprint Aug 2026

HyperODE: Zero-Shot Surrogate for Simulation and Inference of Dynamical Systems

HyperODE is introduced, a surrogate capable of operating across an entire class of approximately mass-conserving compartmental models without retraining, by mapping the structure of ordinary differential equations into directed hypergraphs, which decouples the functional form of system interactions from the neural network architecture.

A. Srivastava · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.