Skip to content
Open access

Data-Driven Modeling and Scientific Computation: Methods for Complex Systems and Big Data

Sep 2026 · Applied Data Science and Analysis · 3 citations

Abstract

The scientific computing field has shifted from using numerical solvers which solved equations to employing hybrid systems which combine physical principles with data-driven modeling techniques. The system shift derives from two interrelated problems because high-dimensional multi-scale nonlinear systems require inaccessible high-performance computing resources to perform detailed discretization and because sensors and simulations and data streams now generate huge amounts of data. The paper evaluates four essential components which form the foundation of contemporary data-based computational tools through Dynamic Mode Decomposition (DMD) and Singular Value Decomposition (SVD)-based Reduced-Order Modeling (ROM) and the Sparse Identification of Nonlinear Dynamics (SINDy) for equation discovery and Physics-Informed Neural Networks (PINNs) which incorporate conservation laws into the training loss function and scalable high-performance-computing (HPC) systems for processing scientific data in real-time. The mathematical frameworks which govern each method receive presentation together with numerical results from original experiments which used four benchmark systems: SINDy-based sparse recovery of the Lorenz-63 chaotic attractor and DMD-based modal decomposition of a synthetic two-frequency spatiotemporal field and POD-based reduced-order modeling of the viscous Burgers equation and PINN-based solution of the 1D heat equation benchmarked against a Crank–Nicolson finite-difference solver. At 5% multiplicative noise, SINDy recovers the governing Lorenz coefficients with a relative  error of 9.1  while preserving exact sparse support, whereas DMD recovers both oscillation frequencies to within 0.05% even at 50% noise despite reconstruction error growing approximately linearly. POD compresses the Burgers solution manifold by 85× at 99% energy capture. Our PINN attains a relative  error of 1.3  with only 161 trainable parameters, while the Crank–Nicolson solver achieves 7.1  in under 9 ms with certified second-order convergence (observed order 2.000) — quantifying precisely the accuracy-per-unit-cost gap that motivates hybrid formulations. We conclude that data-driven methods offer decisive advantages in the many-query, partially-known-physics, and inverse-problem regimes, at the cost of interpretability and rigorous error certification, motivating physics-constrained architectures as the most promising direction forward.  

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.