Skip to content
#data science Dataset Open access

Harmonized Geospatial Dataset from the 2010 Brazilian Demographic Census by IBGE

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Harmonized Geospatial Dataset from the 2010 Brazilian Demographic Census by IBGE This dataset provides a harmonized geospatial version of the 2010 Brazilian Demographic Census at the census tract level. The data are distributed as separate, gzip-compressed files (one per table and per geography layer) together with the script that builds a single GeoPackage from them. The resulting database contains cleaned, standardized, and consistently typed attribute tables linked to official census tract geometries. Variable names were harmonized, data types were explicitly enforced, and structural inconsistencies were resolved to facilitate reproducible spatial analysis and integration with GIS and data science workflows. The original data source is the 2010 Brazilian Demographic Census conducted by the Instituto Brasileiro de Geografia e Estatística (IBGE): the census tract aggregates ("Agregados por Setores Censitários") and the tract synopsis ("Sinopse por Setores"), and the official census tract, district, subdistrict and municipality boundaries. What is in this deposit Total number of files: 34 attribute tables `*.csv.gz` — (basico, domicilio01–02, domiciliorenda, entorno01–05, pessoa01–13, pessoarenda, responsavel01–02, responsavelrenda, sinopse). One row per census tract, first column `setor_code`, then the variables `V001`, `V002`, … geospatial layers `*.gpkg.gz` — setores, subdistritos, distritos, municipios, sedes and bairros. metadata catalogs: `IBGE_CENSO-2010_LAYERS_X_X.csv`, `IBGE_CENSO-2010_FIELDS_X_X.csv`, `IBGE_CENSO-2022_VARIABLES_X_X.csv` — the metadata catalogs (uncompressed, readable on the repository). builder script `build_database.py` — the script that builds a GeoPackage database from the sources files. Files are kept separate on purpose, so that new versions of the deposit only replace the files that changed. The `.csv.gz` files can be read directly by pandas, pyarrow, DuckDB, R and most other tools; the `.gpkg.gz` files must be decompressed before use in GIS software. Rebuilding the database 1. Download all files into one folder (keep the `.gz` files as they are; no need to decompress).2. Create and edit `build_database_2010.json`: set `input_folder` (the folder above), `output_folder` and the three catalog paths.3. Run `python build_database.py --config build_database_2010.json --overwrite`. Requirements: Python 3 (standard library only) and the GDAL command line tool `ogr2ogr` available in the PATH. The script reads the compressed files directly (geography layers are unpacked temporarily next to the output and removed), builds the database from zero, and writes a log with checks (row counts, column types, sector key cross-check). The header of the script documents every option of the configuration. Layers The Database GeoPackage has 34 layers, named without census or year (e.g. `basico`, `setores`): - Geospatial: `setores`, `distritos`, `subdistritos`, `municipios`.- Attribute tables: the 27 tables listed above, linked to `setores` by `setor_code`.- Three structured metadata layers: - `layers` - Describes all database layers contained in the GeoPackage, including geometry type, spatial reference system, thematic scope, and relational structure where applicable. - `fields` - Documents major attribute fields and naming conventions, including standardized prefixes, identifier variables, geographic codes, and structural variables used for relational integrity and joins. - `variables` - Provides a complete listing of original census variables, including official variable names, harmonized names (when applicable), and their corresponding descriptions (Portuguese and English) as defined in the 2010 census documentation. Together, these metadata layers provide full transparency regarding database structure, schema harmonization, and variable provenance. Conventions - Identifiers: every column ending in `_code` (and `code`) is stored as text, so leading zeros and long codes are never altered.- Variables: columns named `V` + number are numeric. They are stored as integers, except the few variables that actually hold decimals, which are detected from the data when the database is built (they are reported in the build log). Counts that the source wrote as `12.0` are stored as whole numbers.- Missing and protected values use fixed codes in every table (documented in the `fields` layer): `-9999` no data (empty in the source), `-1` hidden for statistical confidentiality (`X` in the source), `-2` uncertain (`.` in the source). Text fields use `NA` for no data. Data quality notes The first 2010 harmonized tables had structural defects that were corrected in this version (details in the log and scripts of the repository): - Sector codes that a spreadsheet had turned into scientific notation (e.g. `2,3001E+14`) or into short codes were restored from the order of the household table (`domicilio01`), only where the match was unambiguous. This affected the entorno02–05 tables (72,521 rows), pessoa08 (2,346 rows) and pessoa11 (47,733 rows).- 47,733 rows of São Paulo in responsavel01 and responsavel02 had been written whole inside the sector code cell; they were parsed back into columns.- 47,733 identical duplicate rows in pessoa08 were removed.- Three cells filled with garbage characters (a truncated buffer in the source files: one in entorno05, two in pessoa04) were set to no data (`-9999`).- Empty cells are `-9999` and `X` is `-1` (for example the 888,338 hidden cells of the `sinopse` table). Tables do not all cover exactly the same set of tracts. Most have 297,741 tracts; `basico` has 12 fewer, `pessoa04` 21 fewer, `pessoa07` 28 fewer, entorno05 and responsavel01–02 6 fewer, and `sinopse` has a larger universe (311,052 tracts, 13,317 of them absent from the household table). Join on `setor_code` and expect unmatched rows. Version log v2.0.0 Second version, with the split method: the tables and geography layers are distributed as separate compressed files, with the harmonization fixes above, and users can rebuild the database by running the `build_database.py` script. v1.0.0 First version approach with a single database shipped.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Trajectory Balance: Improved Credit Assignment in GFlowNets

It is proved that any global minimizer of the trajectory balance objective can define a policy that samples exactly from the target distribution, and empirically demonstrate the benefits of the trajectories balance objective for GFlowNet convergence, diversity of generated samples, and robustness to long action sequenc...

Esmeralda S. Whitammer, Moksh Jain, Emmanuel Bengio et al. · 302 citations · ⚡60

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.