Synthetic Food Bank Network Dataset
Abstract
A fully synthetic, openly licensed dataset describing a multi-stakeholder food bank network: clients with demographic attributes and modeled food preferences (stated and revealed), partner pantries with locations and operating hours, grocery-store and residential donors, volunteers with availability and delivery capacity, an inventory catalog, a stream of donation events, simulated pickup, delivery, and volunteer-shift histories, and driving-time origin-destination (OD) matrices connecting block groups, pantries, donors, and clients. The data is modeled on the West Alabama Food Bank nine-county service area but contains no real individuals, addresses, or identifying information. It is generated from open demographic inputs (American Community Survey tables and a fitted spatial point process) together with a preference model fit to a consented research survey. Only fitted model parameters, distributional summaries, and fully regenerated synthetic records are released, so the dataset can be inspected, regenerated, and reused without privacy constraints. Understanding the columns. Start with DATA_DICTIONARY.md in the file list above: every table, column, data type and join key, with an entity-relationship diagram, a join-key reference and cross-stage join recipes. data_dictionary.csv carries the same content as one row per file and column, for tooling, and previews as a table in the browser. README.md orients you to the record. None of the three requires downloading the archive. The dictionary also travels inside the archive at docs/data_dictionary.md, alongside per-stage schema files under 01_synthetic_data_generator/schemas/ and 02_preference_modeling/schemas/. What is in this record. 53 Apache Parquet tables, 6 JSON model-fit diagnostics, 1 Markdown fit summary, and one archive containing the complete project (all of the data together with the generation pipeline, schema documentation, and source code) for users who prefer a single download. Synthetic core entities (12 tables). clients, pantries, pantry_hours, donors_grocery, donors_residential, volunteers, volunteer_availability, inventory_catalog, donations, and the three 12-month behavioral histories client_pickup_history, client_delivery_history, volunteer_shift_history. Spatial backbone (4 tables). wafb_block_groups, wafb_bg_demographics, alabama_counties, ms_block_group_weights. Preferences (6 tables). Per-client synthesized preferences and catalog projections for both the revealed and the stated model, plus the two crosswalks linking mock-store items to catalog stock-keeping units. Fitted preference parameters and diagnostics (4 tables, 4 JSON). Per-item utilities and category-by-demographic coefficients for the revealed and stated models, with their cross-validation and fit diagnostics. Travel-time decay (1 table, 1 JSON). Global and per-tract fitted decay parameters. Pantry utilization, two views (6 tables). The realized view is computed from the simulated pickup history. The predicted view is a Huff-model forecast and is deposited under a huff_ filename prefix, because a Zenodo record is a flat namespace and the two views share base names. They are complementary rather than competing. Pantry-choice model (1 table, 1 JSON, 1 Markdown). The fitted multinomial-logit coefficients and their projection onto the synthetic pantry network. This is the model that generates the pickup history, so it is what you need in order to audit the realized flows or substitute a choice model of your own. Origin-destination travel-time matrices (19 tables). A master location key plus 18 matrices covering six entity-pair configurations across three time-of-day windows. These are the bulk of the volume. How to use the data. The tables are Apache Parquet and load directly in pandas, polars, DuckDB, R (arrow), and Spark. Entities join on client_id, pantry_id, donor_id, volunteer_id, sku_id, and block_group_geoid; the OD matrices join through the location key table. The dictionary's join-key reference gives the full map. Runnable tutorials and worked examples are maintained in a companion GitHub repository that will be made public when the accompanying Data Descriptor is published; until then, the archive in this record is self-contained and carries the pipeline and documentation. How to cite. If you use this dataset, please cite the accompanying Data Descriptor (Nature Scientific Data, in preparation; citation forthcoming). Until it is published, please cite this Zenodo dataset by its DOI. Cite the version DOI to refer to these exact files, or the concept DOI to refer to the dataset generally, in which case the link always resolves to the newest version. License. The dataset and its documentation are released under CC-BY-4.0; the accompanying source code (in the bundled archive and the companion repository) is released under the MIT License. Funded under the U.S. National Science Foundation SHARING program (Award No. 2125600). Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the National Science Foundation.