Skip to content
#software testing Dataset Open access

The Cosmological Constructor: A Backward State-Machine Reconstruction from Measured Data / Космологичен конструктор: възстановяване на състоянията назад от измерени данни

Aug 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

English Dataset and software · bilingual Bulgarian/English release The Cosmological Constructor is an audit-ready dataset and software release centred on append-only, typed scientific memory. The historical FITS files are never used as parents of the new chain and are never rewritten. Their contents are migrated into active semantic memory, while byte-exact snapshots remain available only as immutable provenance evidence. The source history contains 120 append-only FITS states with 274,180 table-row occurrences across 37 table schemas. The active migration stores 94,229 unique typed row values as standard SEMANTIC_OBJECTS. Identical typed values are stored once, but every original occurrence remains addressable by its exact [source sequence, HDU index, row index] position. ASCII, logical and integer cells are explicitly typed; floating-point cells retain their original IEEE-754 bit patterns. No averaging, sampling, binning or numerical rounding is introduced by the migration. The canonical FITS memory consists of nine SHA-256-linked commits. Commits 000002 and 000006 preserve the 120 historical files as byte-exact provenance snapshots. Commits 000007 and 000008 contain the active typed-value migration. The complete typed memory contains 47 action records, 218 layered-memory records, 94,385 semantic objects, 39 conclusions, 102 evidence paths, four cycle summaries and 449 lineage edges. Integrity and reproducibility all 120 source files pass FITS CHECKSUM/DATASUM and source-chain validation; all 274,180 source occurrences match the active typed representation; typed-value verdict: ALL_TYPED_VALUES_AND_OCCURRENCES_EXACT; file-chain verdict: CHAIN_INTEGRITY_OK; JSONL-to-FITS verdict: FULL_JSONL_FITS_OBJECT_EQUIVALENCE_OK; repeating either migration is idempotent and creates no additional commit; the focused FITS, universal-ingest, constant-frame-rate A/V and variable-frame-rate A/V tests pass; all eight runtime modules complete without exceptions, and repeated runtime evidence is byte-identical. Scientific and epistemic scope This revision changes storage, migration, consistency and provenance enforcement; it does not change the physical model, scientific equations, declared thresholds or numerical-precision rules. The machine contract remains fail-closed and additive. Results produced in the supplied runtime are labelled SELF_VALIDATED, INTERNAL_SELF_VALIDATED or INTERNAL_HELDOUT_SELF_VALIDATED. No result is presented as externally VALIDATED; that status requires a genuinely independent source and execution. Package contents 01_CURRENT_SOURCE/ — executable source, migration tools and focused tests; 02_DATA_INPUT/ — frozen, explicitly identified runtime input; 03_MEMORY/ — action ledger, layered semantic memory, cycle history and the nine-commit typed FITS chain; 04_RUNTIME_EVIDENCE/ — deterministic reports and auxiliary ledgers; 09_PROVENANCE_EVIDENCE/ — integrity report and quarantined/superseded artifacts; the release manifest and SHA256SUMS.txt — canonical inventory, sizes and hashes. Verify the downloaded release with sha256sum -c SHA256SUMS.txt. Python requires NumPy, Astropy and mpmath; the audio/video focused tests additionally require FFmpeg. Source-specific provenance and licensing notes remain part of the release documentation. Български Данни и софтуер · двуезично издание на български и английски Космологичният конструктор е одитируемо издание на данни и софтуер, изградено около append-only типизирана научна памет. Историческите FITS файлове никога не са родители на новата верига и никога не се презаписват. Съдържанието им е мигрирано в активната семантична памет, а byte-exact snapshots са запазени единствено като неизменимо доказателство за произхода. Изходната история съдържа 120 append-only FITS състояния с 274 180 срещания на таблични редове в 37 таблични схеми. Активната миграция съхранява 94 229 уникални типизирани стойности на редове като стандартни SEMANTIC_OBJECTS. Еднаквите типизирани стойности се пазят веднъж, но всяко първоначално срещане остава адресируемо чрез точната позиция [пореден номер на източника, HDU индекс, индекс на реда]. ASCII, логическите и целочислените клетки са изрично типизирани; клетките с плаваща запетая пазят първоначалните си IEEE-754 битове. Миграцията не въвежда усредняване, sampling, binning или числово закръгляне. Каноничната FITS памет се състои от девет SHA-256-свързани commits. Commits 000002 и 000006 пазят 120-те исторически файла като byte-exact provenance snapshots. Commits 000007 и 000008 съдържат активната миграция на типизираните стойности. Пълната типизирана памет съдържа 47 action записа, 218 layered-memory записа, 94 385 семантични обекта, 39 заключения, 102 доказателствени пътя, четири cycle summaries и 449 lineage връзки. Цялост и възпроизводимост всичките 120 изходни файла преминават FITS CHECKSUM/DATASUM и проверката на изходната верига; всичките 274 180 изходни срещания съвпадат с активното типизирано представяне; присъда за типизираните стойности: ALL_TYPED_VALUES_AND_OCCURRENCES_EXACT; присъда за файловата верига: CHAIN_INTEGRITY_OK; присъда за JSONL-to-FITS: FULL_JSONL_FITS_OBJECT_EQUIVALENCE_OK; повторното изпълнение на всяка миграция е идемпотентно и не създава допълнителен commit; focused тестовете за FITS, universal ingest, constant-frame-rate A/V и variable-frame-rate A/V преминават; всичките осем runtime модула завършват без изключения, а повторните runtime evidence файлове са byte-identical. Научен и епистемичен обхват Тази ревизия променя съхранението, миграцията, консистентността и контрола на произхода; тя не променя физическия модел, научните уравнения, обявените прагове или правилата за числова точност. Машинният договор остава fail-closed и additive. Резултатите от приложения runtime са означени като SELF_VALIDATED, INTERNAL_SELF_VALIDATED или INTERNAL_HELDOUT_SELF_VALIDATED. Нито един резултат не е представен като външно VALIDATED; този статус изисква действително независим източник и независимо изпълнение. Съдържание на пакета 01_CURRENT_SOURCE/ — изпълним source, инструменти за миграция и focused тестове; 02_DATA_INPUT/ — замразен и еднозначно идентифициран runtime вход; 03_MEMORY/ — action ledger, layered semantic memory, cycle history и деветкомитова typed FITS верига; 04_RUNTIME_EVIDENCE/ — детерминистични отчети и помощни ledgers; 09_PROVENANCE_EVIDENCE/ — integrity report и карантинни/заменени артефакти; manifest файлът на изданието и SHA256SUMS.txt — каноничен опис, размери и хешове. Проверката на сваленото издание се изпълнява с sha256sum -c SHA256SUMS.txt. Python средата изисква NumPy, Astropy и mpmath; focused тестовете за аудио и видео изискват допълнително FFmpeg. Бележките за произхода и лицензите на отделните източници остават част от документацията на изданието.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#computer vision Review Mar 2008

Agile methods in European embedded software development organisations: a survey on the actual use and usefulness of Extreme Programming and Scrum

The results show that the embedded industry has been able to apply agile methods in its development processes and that the appreciation of the agile methods and their individual practices appears to increase once adopted and applied in practice.

O. Salo, P. Abrahamsson · 238 citations · ⚡9
#computer vision Open access Jul 2017

What happens when software developers are (un)happy

Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 236 citations · ⚡13
#computer vision Open access Oct 2004

Mobile-D: an agile approach for mobile application development

The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.

P. Abrahamsson, Antti Hanhineva, H. Hulkko et al. · 225 citations · ⚡18

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.