Farm typology and livelihood vulnerability of 421 farms in the upper río Hacha basin, Caquetá, Colombia
Abstract
De-identified household-survey data for 421 farms in the upper río Hacha basin (Florencia, Caquetá, Colombia), surveyed December 2022 to March 2023 under a Participatory Rural Appraisal. How this deposit is organisedFiles 01 to 06 are the reusable base: the farm typology, the livelihood vulnerability indices, the exposure climate metrics, the adaptive-capacity indicator metadata, the spatial weights matrix and the continuous outcome variables in grouped form. File 00 is the variable dictionary for all files. What changed in version 2The exposure index was rebuilt from fine-resolution climate data following Hahn et al. (2009): the interannual standard deviation (1992-2022) of monthly precipitation (CHIRPS v2.0, 0.05°) and of the monthly mean of daily maximum and minimum temperature (ERA5-Land, 0.1°), averaged over the 12 calendar months, interpolated to each farm, min-max normalised and averaged with equal weights. It replaces the version 1 index (principal component analysis of twelve NASA POWER metrics), which took only five distinct values across the basin. Files 00, 02, 03 and 05 were updated; files 01, 04 and 06 are unchanged. Contents- Farm typology built from 42 active categorical variables through Multiple Correspondence Analysis with Benzécri correction, Ward.D2 hierarchical clustering and k-means consolidation (k = 3; 5 dimensions retained, 81.1% corrected inertia; silhouette 0.248; bootstrap ARI 0.74). The three types are Smallholder subsistence cropping (n = 257), Consolidated mixed farming (n = 107) and Low-activity holdings (n = 57).- Livelihood vulnerability assessment: exposure (interannual variability of precipitation and of maximum and minimum temperature, 1992-2022), sensitivity and its four subcomponents, adaptive capacity in a full and a robustness specification, and the composite index (IVL = IExpo + ISes - ICadp, min-max rescaled to 0-1 and grouped into five equal-interval classes). Privacy protectionDirect identifiers and farm coordinates are removed, survey identifiers replaced by pseudonyms, and veredas with fewer than five farms pooled. Continuous outcome variables are published in grouped form (age in 5-year bands, area in 10 ha steps). Because the new exposure index varies continuously in space, the farm-level exposure index is released rounded to two decimals, the indices derived from it are computed from the rounded value, and the three exposure indicators are released only as vereda means (file 03). In place of coordinates, the k-nearest-neighbour spatial weights matrix used in the spatial analysis (the 380 farms inside the basin boundary; k = 8, row-standardised) is released as a pseudonymised edge list, with rows in the order used in the analysis. Recomputing from these files returns Moran's I = 0.52 (p = 0.001) for the vulnerability index and the published LISA cluster counts within two farms; the small differences from the article (Moran's I = 0.521) arise from the rounding of exposure.