World Cuisine Ingredient Overlap (26 Cuisines, 7 Languages)
Ingredient overlap between 26 world cuisines, measured over 1,956 well-known dishes described in 7 languages. Provenance — read this first: the dishes are real, well-known dishes, but their ingredient lists were written by a large language model and verified by a second model pass, not transcribed from cookbooks. This dataset therefore measures how a language model represents 26 cuisines, under a single consistent describer, rather than how people actually cook. A manual audit found ingredient/dish mismatches in about 7% of entries. It should not be cited as an ethnographic record. Four tables: 26 cuisines with dish counts and country centroids; 958 canonical ingredients with names in 7 languages; a bipartite cuisine-to-ingredient graph of 3,164 rows; and all 289 cuisine pairs with overlap share, great-circle centroid distance and a land-border flag. The distance and border columns support the obvious test. Across all pairs, Pearson r between overlap and distance is -0.33, and -0.28 after removing every pair sharing a land border. The 14 pairs that share a land border average 2.01% overlap against 0.63% for the 275 that do not, and the closest quarter of pairs averages 1.26% against 0.39% for the farthest quarter. A handful of pairs resist this, and the strongest link in the corpus is one of them: Ethiopia-Philippines at 8,993 km, ahead of every land-border pair. Mexico-Philippines lands at 13,844 km and UK-US at 6,952 km. This is not an artefact of corpus size — size correlates with overlap at only -0.18, and Vietnam (52 dishes, 0.40% mean overlap) sits at the bottom while Mexico (52 dishes, 0.85%) sits near the top. Version 2.0.0 adds Russian and Ethiopian cuisines (24 to 26) and reflects a duplicate-removal pass on the corpus.