Skip to content
#large language models Dataset Open access

Reproducibility Package for "A Governance Architecture for AI-Assisted Validation of Construction Cost Knowledge: Combining Domain Rules with LLM Semantic Analysis"

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This deposit contains the reproducibility materials for the manuscript "A Governance Architecture for AI-Assisted Validation of Construction Cost Knowledge: Combining Domain Rules with LLM Semantic Analysis." The manuscript presents a deployed tender management system that validates Bill of Quantities (BOQ) pricing formulas at portfolio scale by combining trade-specific deterministic policy checks with large language model (LLM) semantic analysis under a four-layer hallucination-mitigation architecture (schema enforcement, evidence binding, deterministic override, and read-only AI). The evaluation includes 17 construction tenders across Morocco and Ivory Coast, representing approximately 1.1 billion monetary units of portfolio value, and 100 expert-labelled formulas with an inter-rater agreement of Cohen's κ = 0.96. Package contents (10 files): README.docx — Package overview, file listing, sanitization policy, and citation guidance. CHANGELOG.txt — Version history describing updates to each release. Reproducibility_Supplement.pdf — Dataset definitions, key statistics, leakage-prevention protocol, baseline model specifications, statistical analysis details, the complete verbatim LLM prompt template, and a worked input-output JSON trace for the REV-002 example. LLM_Prompt_Template_Full.docx — Complete verbatim system and user-message templates used with the Azure OpenAI gpt-5.1-codex-mini deployment. Results.xlsx — Headline manuscript results, including Expert Metrics, Confusion Matrices, Cross-Region F1, Risk Distribution, and the Data Dictionary. scripts/verify_from_deposit.py — Verification script that recomputes the manuscript's expert-sample confusion matrices, overall and per-region F1 scores, and portfolio risk distribution directly from the deposited data, printing the recomputed and reported values side by side. Data/Expert_Validation_Sample.xlsx — The 100 expert-labelled formula records, including the associated construction trade for each formula. Data/Validation_Mode_Predictions.xlsx — Per-formula predictions for the four validation configurations (Hybrid, Rules-only, AI-only, and AI-with-flags). Data/Baseline_Predictions.xlsx — Outputs for the six baseline models (LR-S2, XGB-S2, RF-S2, Isolation Forest, z-score, and TF-IDF+LR), together with the baseline ranking and fold composition. Data/Statistical_Tests_Frozen.xlsx — Official statistical-test artefacts, including McNemar's test, i.i.d. and TenderId-cluster bootstrap confidence intervals, per-fold stability, and inter-rater agreement. Sanitization. Item descriptions, unit prices, and AI-generated narrative text have been removed, while FormulaId and TenderId have been pseudonymized. Component percentages, construction trades, expert labels, and risk classifications—the variables used in the manuscript's analyses—remain intact. Consequently, all quantitative results reported in the manuscript can be reproduced from the released artefacts. Version notes. This release aligns per-formula predictions across validation modes using the composite key (TenderId, FormulaId), adds the per-formula Trade column, and includes the verification script. See CHANGELOG.txt for the complete revision history. Raw BOQ text and original monetary values are not publicly released because they contain commercially confidential contractor pricing from active projects. Researchers requiring access to additional anonymized data may contact the corresponding author. Version 5 (September 2026) accompanies the revised manuscript. It adds the trade-label audit workbook (Data/Trade_Label_Audit.xlsx, 146 screened candidates and 80 controls with the senior estimator's verdicts) and its screening script, the three-model replication outputs (Data/second-model/, per formula per arm per configuration), a CHANGELOG.txt describing every change since v4, and corrected Region and component columns in the expert-sample workbook; no reported metric changed. README.docx lists the contents; scripts/verify_from_deposit.py recomputes every headline number from the deposited files.

View source

Similar papers

#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Conference Open access Dec 2013

Affordable and Energy-Efficient Cloud Computing Clusters: The Bolzano Raspberry Pi Cloud Cluster Experiment

The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.

P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al. · 110 citations · ⚡7
#computer vision Book Open access Mar 2017

On the Unhappiness of Software Developers

The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 84 citations · ⚡6

Related blog posts

MIT News · Artificial Intelligence Sep 14, 2026

New method enables AI for safety-critical situations

The “HardFlow” algorithm could help generative AI models produce high-quality outputs that obey strict requirements when “pretty close” doesn’t cut it.

GPT-Lab Sep 10, 2026

Responsible AI Must Consider Its Afterlife

AI may appear weightless, but every model depends on physical infrastructure. To understand responsible AI, we need to look beyond algorithms and consider the entire lifecycle of the hardware behind them. The post Responsible AI Must Consider Its Afterlife appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.