Harmonized Geospatial Dataset from the 2010 Brazilian Demographic Census by IBGE
Abstract
Harmonized Geospatial Dataset from the 2010 Brazilian Demographic Census by IBGE This dataset provides a harmonized geospatial version of the 2010 Brazilian Demographic Census at the census tract level. The data are distributed as separate, gzip-compressed files (one per table and per geography layer) together with the script that builds a single GeoPackage from them. The resulting database contains cleaned, standardized, and consistently typed attribute tables linked to official census tract geometries. Variable names were harmonized, data types were explicitly enforced, and structural inconsistencies were resolved to facilitate reproducible spatial analysis and integration with GIS and data science workflows. The original data source is the 2010 Brazilian Demographic Census conducted by the Instituto Brasileiro de Geografia e Estatística (IBGE): the census tract aggregates ("Agregados por Setores Censitários") and the tract synopsis ("Sinopse por Setores"), and the official census tract, district, subdistrict and municipality boundaries. What is in this deposit Total number of files: 34 attribute tables `*.csv.gz` — (basico, domicilio01–02, domiciliorenda, entorno01–05, pessoa01–13, pessoarenda, responsavel01–02, responsavelrenda, sinopse). One row per census tract, first column `setor_code`, then the variables `V001`, `V002`, … geospatial layers `*.gpkg.gz` — setores, subdistritos, distritos, municipios, sedes and bairros. metadata catalogs: `IBGE_CENSO-2010_LAYERS_X_X.csv`, `IBGE_CENSO-2010_FIELDS_X_X.csv`, `IBGE_CENSO-2022_VARIABLES_X_X.csv` — the metadata catalogs (uncompressed, readable on the repository). builder script `build_database.py` — the script that builds a GeoPackage database from the sources files. Files are kept separate on purpose, so that new versions of the deposit only replace the files that changed. The `.csv.gz` files can be read directly by pandas, pyarrow, DuckDB, R and most other tools; the `.gpkg.gz` files must be decompressed before use in GIS software. Rebuilding the database 1. Download all files into one folder (keep the `.gz` files as they are; no need to decompress).2. Create and edit `build_database_2010.json`: set `input_folder` (the folder above), `output_folder` and the three catalog paths.3. Run `python build_database.py --config build_database_2010.json --overwrite`. Requirements: Python 3 (standard library only) and the GDAL command line tool `ogr2ogr` available in the PATH. The script reads the compressed files directly (geography layers are unpacked temporarily next to the output and removed), builds the database from zero, and writes a log with checks (row counts, column types, sector key cross-check). The header of the script documents every option of the configuration. Layers The Database GeoPackage has 34 layers, named without census or year (e.g. `basico`, `setores`): - Geospatial: `setores`, `distritos`, `subdistritos`, `municipios`.- Attribute tables: the 27 tables listed above, linked to `setores` by `setor_code`.- Three structured metadata layers: - `layers` - Describes all database layers contained in the GeoPackage, including geometry type, spatial reference system, thematic scope, and relational structure where applicable. - `fields` - Documents major attribute fields and naming conventions, including standardized prefixes, identifier variables, geographic codes, and structural variables used for relational integrity and joins. - `variables` - Provides a complete listing of original census variables, including official variable names, harmonized names (when applicable), and their corresponding descriptions (Portuguese and English) as defined in the 2010 census documentation. Together, these metadata layers provide full transparency regarding database structure, schema harmonization, and variable provenance. Conventions - Identifiers: every column ending in `_code` (and `code`) is stored as text, so leading zeros and long codes are never altered.- Variables: columns named `V` + number are numeric. They are stored as integers, except the few variables that actually hold decimals, which are detected from the data when the database is built (they are reported in the build log). Counts that the source wrote as `12.0` are stored as whole numbers.- Missing and protected values use fixed codes in every table (documented in the `fields` layer): `-9999` no data (empty in the source), `-1` hidden for statistical confidentiality (`X` in the source), `-2` uncertain (`.` in the source). Text fields use `NA` for no data. Data quality notes The first 2010 harmonized tables had structural defects that were corrected in this version (details in the log and scripts of the repository): - Sector codes that a spreadsheet had turned into scientific notation (e.g. `2,3001E+14`) or into short codes were restored from the order of the household table (`domicilio01`), only where the match was unambiguous. This affected the entorno02–05 tables (72,521 rows), pessoa08 (2,346 rows) and pessoa11 (47,733 rows).- 47,733 rows of São Paulo in responsavel01 and responsavel02 had been written whole inside the sector code cell; they were parsed back into columns.- 47,733 identical duplicate rows in pessoa08 were removed.- Three cells filled with garbage characters (a truncated buffer in the source files: one in entorno05, two in pessoa04) were set to no data (`-9999`).- Empty cells are `-9999` and `X` is `-1` (for example the 888,338 hidden cells of the `sinopse` table). Tables do not all cover exactly the same set of tracts. Most have 297,741 tracts; `basico` has 12 fewer, `pessoa04` 21 fewer, `pessoa07` 28 fewer, entorno05 and responsavel01–02 6 fewer, and `sinopse` has a larger universe (311,052 tracts, 13,317 of them absent from the household table). Join on `setor_code` and expect unmatched rows. Version log v2.0.0 Second version, with the split method: the tables and geography layers are distributed as separate compressed files, with the harmonization fixes above, and users can rebuild the database by running the `build_database.py` script. v1.0.0 First version approach with a single database shipped.