Skip to content
#software testing Open access

closure_drift: does your version label name exactly one version of your code?

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research) · 1 citation
Software Engineering Research

Abstract

A read-only, zero-dependency tool that measures, over any git repository, whether a declared version label identifies exactly one state of the producing code at the points where that repository publishes. Addressing a published artefact by (input, version) is sound only if the label is injective over closures. Nothing enforces that: the label is a string a human edits. When two code states share a label, one address denotes two outputs, and the system cannot detect it, because the label is the only thing it recorded. The tool distinguishes publication regimes. At tags (the default) it measures the case of released software; at commits it measures continuously published output. Measured against four widely used open-source projects (click, requests, packaging, httpx) it reports no drift; against a system publishing a daily edition under a hand-maintained label it reports one label covering six distinct closures. The failure belongs to continuous publication, not to versioning in general. Each report stamps the commit measured and the hash of the tool that measured it, because a count over repository history is a function of repository state. Version 0.4.0: the measurement script is byte-identical to 0.3.0; this version adds a negative test fixture (label-only ledger: must be refused at tags and reported as drift at commits), the documented limitation (the detector checks addressing, not re-execution of originating code states), and new reference results over three public repositories selected under a pre-registered rule. Version 0.5.0: the measurement script is byte-identical to 0.3.0 and 0.4.0, and no reference result was re-run. What changes is what the deposit says about itself, and what it asks for. RESULTS.md is a table for measurements produced by someone other than the author, on repositories the author does not control; it is published empty, because as of this release nobody outside the author has run the tool and reported a result, and omitting the section would let a reader assume otherwise. It states what a line must carry to count — the version DOI of the deposit used, the stamp block as emitted, and the publication-point setting — and records that results contradicting the detector are wanted on the same terms as results confirming it. The report carries counts, labels and hashes and never file contents, so a private repository can be measured without anything leaving the machine. SCOPE.md states what the tool does and, explicitly, what it will not be extended to do: the (A) the record is well-formed / (B) the artefact can be re-produced boundary, restated as a commitment rather than a caveat. NOTICE records that the author has patent applications pending; it adds no condition to the licence, and commercial use carries no royalty and no payment obligation. The README now leads with the measured result and adds two sections: why an unambiguous address is a precondition of reproducibility rather than a part of it, and what this tool is not — it never rebuilds and it issues no attestation. Version 0.6.0: the measurement script is byte-identical to 0.3.0, 0.4.0 and 0.5.0 (sha256 da5da3c0e781b67b9b3a55800d599c243edc8df649fc90b24a88e289533805c5), so every result produced under any of those deposits remains valid and comparable, and no measurement was re-run. This version corrects the licence file and publishes the source. LICENSE now carries the unabridged text of the Apache License, Version 2.0. The file deposited as 0.3.0 through 0.5.0 was an abridged text of that licence — 1,064 words against 1,581 — with the copyright grant, the patent grant and the trademarks section word-for-word identical, but the definitions, redistribution, warranty-disclaimer and liability-limitation sections shortened and the appendix absent. The licence named in every deposit has always been Apache-2.0 and the grant sections a licensee relies on were already canonical; from this version the file says so too, and the changelog records exactly what was missing. A record that looks well-formed and is not, in this project's own deposit, is the phenomenon this tool exists to detect, and it is recorded rather than quietly replaced. The source repository is now public at https://github.com/luizfnsilva/closure_drift, where the nine files of this deposit are held byte-identical to it under a checksum gate checked on every push, so what is run can be verified without trusting the author. Measurements can now be reported through an issue form as well as by email, and so can the more valuable case — that the detector is wrong about a repository. RESULTS.md carries both channels and remains published empty until a third party sends a line. Version 0.7.0: the measurement script changes, for the first time since 0.3.0. It is sha256 6d8906ef374b73e6b8c58adba813c77c4ff352f5c9c280aa43ff2baa4f804451, against da5da3c0e781b67b9b3a55800d599c243edc8df649fc90b24a88e289533805c5 for 0.3.0 through 0.6.0. What changed is a refusal and not a measurement: a --version-regex that is malformed, or that has no capture group, used to raise an exception and leave the process at exit 1, which is this tool's code for drift — in a detector, an error that cannot be told apart from a finding is the worst possible outcome. Both now refuse by name at exit 2. Under a well-formed configuration no completed measurement changes, and this was checked rather than asserted: the two versions were run over the same repositories and every verdict field is identical, the one field that differs being the detector's own stamp. One kind of run does change, and it is the point of the release: a group-less pattern that never matched used to complete at exit 0 with verdict no_labels, which reads as "this repository declares no version label" when what was broken was the pattern. 0.7.0 was tagged as source and never deposited; 0.7.1 supersedes it. Version 0.7.1: the measurement script is byte-identical to 0.7.0. This version corrects the deposit's description of itself. Three of the nine files still described 0.6.0 after the 0.7.0 release — the citation file still carried version 0.6.0, and the README and SCOPE.md still printed the superseded hash as the script's identity. In a tool whose question is whether a version label names exactly one state of the code, a deposit labelled 0.7.0 whose citation file says 0.6.0 is the failure this tool exists to detect, occurring in its own packaging; it is corrected in a new version rather than by moving the published 0.7.0 tag, which would make that label name two states. The deposited fixture now finds the detector deposited beside it. It resolved the detector one directory up — correct in the source repository, where it lives in tests/, and wrong in this deposit, where the files are flat: unpacked on its own it ended in a traceback, and unpacked beside a directory holding a different closure_drift.py it printed "fixture ok" having measured that other file. It now looks beside itself first, names the detector it ran and the first 16 hex of its sha256 in every outcome, and refuses when none is found. The README now states the exit codes as a closed set and records three measured conditions that were written nowhere, including that a negative --max-commits silently narrows the scanned range. One correction to the paragraph above: the checksum gate checked on every push compares the nine files to the manifest that ships with them, offline — a package cannot establish its own provenance — and the comparison against the published record is a separate, scheduled job. Version 0.9.0: the detector is corrected and the measurement script changes (sha256 6548f891a826034c35ef83b276578c79564b57f892a54422af5c1be44591137c). 0.7.1 was audited with pre-registered proofs, mutation controls and three adversarial passes; against 0.7.1 as deposited, 30 of 51 proofs were red. Fixed: files with non-ASCII names and submodule pointers were silently outside the closure; root-level tests/, docs/ and *.md were inside it; undetermined results exited 0; measuring a repository could run commands named in its git configuration; and the version was read from one file chosen at HEAD, so most older tags were never compared. The label is now found at each tag where the project keeps it. Exit 0 means clean and nothing else. Added: a check to run before tagging (--would-tag), tag families (--tags), --strict, --explain, --compare, --diagnose and a documented JSON report format. A study of the 100 most-downloaded PyPI projects, protocol fixed before the first run and both runs kept, is in the source repository: in 29 of the 92 repositories where a determination was reached, a version label names more than one code state at a tag. 0.9.0 compares more tags than 0.7.1, so earlier results should be measured again. There is no 0.8.0: it was this release's working name and was never tagged or deposited. Version 0.9.1: evidence, and one fix it led to (script sha256 89b5349928eba22b0394d01940ed3d4aa989d6820f48fdddf189ad521689c51c). A property suite checks 17 relations over 60 generated repositories against an independent oracle, written from a specification by a reviewer who did not read the detector; ten controls show each property can fail. It found one defect: on Windows closure globs were matched without regard to case, so one commit could have two closures. That is corrected; on Linux and macOS no report changes. Ten large repositories were measured under a protocol written first (the Linux kernel, LLVM, CPython and seven more): no run exceeded the limit, every tag of the kernel takes 90 seconds and 3.4 GB, and the oracle agrees at every pair of tags with files. The same benchmark shows where the defaults fail, and these are published as open: exit 0 with 9 of 5,508 tags compared unless --strict is passed, a version read from the wrong file on git/git and on DefinitelyTyped, and no version label found in six of ten. Added: THREAT_MODEL.md, docs/FAILURES.md, reproduce.sh. The study

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#computer vision Review Mar 2008

Agile methods in European embedded software development organisations: a survey on the actual use and usefulness of Extreme Programming and Scrum

The results show that the embedded industry has been able to apply agile methods in its development processes and that the appreciation of the agile methods and their individual practices appears to increase once adopted and applied in practice.

O. Salo, P. Abrahamsson · 238 citations · ⚡9
#computer vision Open access Jul 2017

What happens when software developers are (un)happy

Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 236 citations · ⚡13
#computer vision Open access Oct 2004

Mobile-D: an agile approach for mobile application development

The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.

P. Abrahamsson, Antti Hanhineva, H. Hulkko et al. · 225 citations · ⚡18

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.