Research Data for “From Extraction Accuracy to Analytical Admissibility: A Human–AI Division of Labor Protocol for LLM-Based Supply-Chain Risk Intelligence”
This dataset contains the reproducibility materials supporting the study “From Extraction Accuracy to Analytical Admissibility: A Human–AI Division of Labor Protocol for LLM-Based Supply-Chain Risk Intelligence.” The package supports an offline benchmark of four open-weight large language models—DeepSeek-V3.2, Qwen3.8-Max, Kimi K2.5, and GLM-5.3—for supply-chain risk intelligence under a text-bounded evaluation protocol. It includes two datasets: an internal development set of 58 English-language nickel-supply articles and a held-out general-news set of 97 articles, including 11 risk-positive cases. Model inputs are restricted to four fields: article identifier, headline, publication date, and excerpt/summary. The archived materials include frozen labels, prompt architectures, public model outputs, scoring scripts, result tables, data dictionaries, and documentation. The evaluated tasks cover risk-relevance screening, Tier-1 and Tier-2 mechanism classification, temporal judgment using TRUE/FALSE/UNKNOWN labels, and an operational human-review triage analysis. The H032 case is retained as an illustrative definitional-boundary case. The package reproduces the manuscript’s main screening, mechanism-classification, temporal, and triage results from archived model outputs without rerunning any model API. Full copyrighted news articles, private reasoning traces, credentials, personal information, and local machine paths are not included. The dataset is intended to support transparent verification of the reported human–AI division-of-labor analysis.
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.
P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al.· IEEE International Conferenc...· 110 citations· ⚡7
The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· International Conference on...· 84 citations· ⚡6
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 8, 2026