GeoID-PINN, a physics-informed neural network (PINN) for susceptible-infectious-recovered-deceased (SIRD) dynamics is introduced and similar performance across plausible priors supports structured regularization but not unique edge recovery.
Abstract
Regional surveillance data reflect local transmission, reporting, seeding, and external infection pressure, which are difficult to identify separately. We introduce GeoID-PINN, a physics-informed neural network (PINN) for susceptible-infectious-recovered-deceased (SIRD) dynamics. The model represents spatial dependence with a row-stochastic source-composition matrix whose rows assign nonnegative source weights that sum to one. We regularize this matrix toward a spatial prior constructed from distance, adjacency, commuting, or lead-lag information. In a four-region simulation with known truth, a compatible distance prior gives source-composition error 0.099. The error rises to 0.159 without regularization and 0.577 under a strongly misspecified prior, while trajectory fit and transmission-scale estimates remain similar. Accurate trajectories therefore do not guarantee recovery of the regional dependence structure. We also evaluate GeoID-PINN retrospectively using COVID-19 data from 64 Louisiana counties. Relative to an autoregressive negative-binomial baseline, Forecast-Trained Geo-PINN reduces mean squared error (MSE) from 32,957 to 11,468 and mean absolute error (MAE) from 70.60 to 57.73. The baseline has lower negative log likelihood (NLL), 5.158 versus 5.346, indicating better distributional fit but worse point accuracy. In a controlled 15-county comparison, county adjacency reduces MSE by 6.85 percent and MAE by 3.1 percent. Similar performance across plausible priors supports structured regularization but not unique edge recovery. These results require prior-sensitivity and observation-model checks before interpretation.
This paper pairing Failure-Informed PINNs with a Self-Adaptive Importance Sampling (SAIS) refinement strategy, and applying the combination to a nine-equation Susceptible/Vaccinated/Exposed/Infected/Recovered (SVEIR) model that follows three viral strains together with a vaccination compartment, shows an improvement of two to three orders of magnitude.
Kawtar Idhammou Ouyoussef, J. El Karkri, L. M. Tine et al.· Discover Applied Sciences· 0 citations
Public-health surveillance systems rely on downstream indicators to infer latent infection incidence, but delays and observation noise provide only an indirect and temporally distorted view of the underlying epidemic process. Reconstructing upstream epidemic trajectories from these observations is therefore an ill-posed inverse problem, in which different reconstruction assumptions may produce different trajectories that remain consistent with the observed data. Here, we develop a general spectral framework that quantifies the statistical distinguishability of candidate upstream trajectories under delayed and noisy observations. We show that epidemiological delay distributions impose a frequency-dependent temporal resolution limit on epidemic surveillance, fundamentally constraining the distinguishability of rapid upstream variation. This limitation propagates to epidemiological inference, making some quantities substantially more sensitive to reconstruction assumptions than others and rendering distinct event-impact profiles difficult to distinguish from downstream observations.
J. Tang, B. Wilder, R. Rosenfeld· medRxiv· 0 citations
Infectious disease time series are often used to estimate a pathogen's basic reproduction number, $R_0$. However, fits of epidemic models to time series conflate pathogen transmissibility with pre-existing population immunity, so only the *effective* reproduction number, $R_{eff}$, can be inferred. This composite parameter is the product of the underlying $R_0$ and the pre-epidemic susceptible fraction, $x^-$. We show that a conservation law associated with epidemic momentum---prevalence weighted by potential to infect---makes it possible to disentangle transmissibility from prior immunity and to infer $R_0$ and $x^-$ separately from a single epidemic time series. We test the methodology using stochastic epidemic simulations, and illustrate the approach with a reappraisal of influenza transmissibility during the 1918 pandemic, estimating rather than assuming the degree of prior population immunity. For the autumn wave in Philadelphia, USA, we find $R_0\approx2.7$ and $x^-\approx0.8$, implying that about 20\% of the population was already immune before that wave, plausibly as a result of infection during the spring 1918 herald wave.
Identifying spatial origins of biological invasions, disease outbreaks, or environmental contaminants is critical for timely intervention. However, existing methods struggle to resolve overlapping signals from multiple sources or account for extreme zero/one inflation in bounded data. We developed HiBASIL (Hierarchical BAyesian Source Inference and Localization), a Bayesian framework that jointly infers source coordinates and mechanistic dispersal kernels from zero-one-inflated spatial observations. We systematically tested HiBASIL across 2,500 simulations spanning localized to long-distance dispersal regimes, various foci weight mixtures (0.5/0.5 to 0.95/0.05), and varying sample sizes (N = 50 to 500) to evaluate geometric sensitivity, spatial robustness, parameter recovery capability, and multi-focal resolution. We applied HiBASIL in tracking cucurbit downy mildew dispersal in field experiments including non-inoculated control, single-focus, and two-foci treatments arranged in a randomized complete block design. To demonstrate broad applicability, we validated HiBASIL with historical London cholera epidemic data to determine whether it could accurately localize the epicenter. HiBASIL achieved 100% accuracy in kernel identification under correct model specification, 100% and 98% (2% predictive equivalence) in under- and over-parameterization scenarios, and maintained appropriate parsimony in null scenarios. Source localization was highly precise, with a median bias of 0.11 m, successfully resolving minority sources contributing only 5% of the observations. Importantly, the framework provides intrinsic self-diagnostics, signaling model complexity mismatch through posterior bimodality (under-parameterization) or parameter collapse (over-parameterization). HiBASIL is applicable in various epidemic scenarios, including diffused long-distance dispersal. Under an isotropic process, HiBASIL achieved sub-meter mean errors on two-source localization in cucurbit downy mildew field tests, effectively isolating transmission signals from landscape noise despite sparse sampling. Despite confounding influence of underlying network processes in the historical cholera data, the posterior localized the epicenter to within ∼ 33m, demonstrating portability beyond plant disease epidemiology. In summary, HiBASIL can accurately localize multiple sources across various epidemic scenarios, extremely unbalanced foci mixtures, and sparse sampling conditions. HiBASIL also demonstrates wide applicability across agricultural and human disease epidemic systems. By providing an open-source Python implementation, HiBASIL enables rigorous inverse spatial inference for ecology, epidemiology, and environmental monitoring, transforming how discrete transmission sources are identified in complex landscapes.
Fangfang Guo, Sharmodeep Bhattacharyya, S. Chatterjee et al.· bioRxiv· 0 citations
This work proposes MechGNN-Epi, a hybrid framework that couples a spatiotemporal graph encoder with a differentiable SIR update that yields epidemiologically constrained trajectories and produces region- and time-indexed parameter proxies that can be inspected as diagnostic signals, while not being guaranteed as causally identifiable mechanistic parameters.
Given that a disease case is known to have occurred in a specific county or state, can one reconstruct a precise geographic observation for that case? The answer is that one cannot. We instead sample multiple high-probability locations for that case under the assumption that cases are more likely to occur in places where there are more people. To do this, we develop a novel and user-friendly R package called USPopulationSampler that randomly samples geospatial locations within census block groups (BG; the smallest geographic unit for which population counts are available) within target counties, states, or across the entirety of the U.S., randomly selected according to population counts using Census Bureau reference data. We use the package to generate 28 realistic high-probability data replicates of 100M+ synthetic spatiotemporal Covid-19 cases observed between the dates of January 21st, 2020 to March 23rd, 2023 and openly publish these data replicates on Zenodo for easy access. Given the large scale nature of the data, the USPopulationSampler package also provides tools for fast download of these datasets and functions to generate further replicates at scale using multi-core parallelization.
R. T. Yean, J. Zhang, A. J. Holbrook· medRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.