Reproducibility artifact for the manuscript Evaluating LLM-Based Credit Rating Systems: A Pre-Registered Specification-Curve Analysis. A pre-registered specification-curve (multiverse) study of large-language-model credit-rating instability. From a frozen corpus of 8,640 elicitations (90-item firm-profile battery with an objective Altman Z'' benchmark, crossed with a 32-specification factorial grid of elicitation choices spanning provider, model version, temperature, prompt paraphrase, output format, few-shot exemplars and answer presentation, three seeds, two providers) the analysis regenerates the OLS (Type-II ANOVA) variance decomposition, the honored-determinism permutation test, within/cross-vendor Fleiss-kappa agreement, the granularity-kappa ladder, and the economic translation (portfolio turnover, Cornaggia-anchored spread, a post-hoc Basel CRE20 capital recast, and the RCAP calibration). A decontaminated real-firm arm (31 anonymized, per-issuer-perturbed issuers benchmarked to disclosed agency ratings, under a pre-registered fingerprinting gate) shows the WATCH-stratum flip-share replicates (consistent with the constructed battery; the difference is not distinguishable from zero, though a formal TOST equivalence test is inconclusive at this sample size). A post-hoc flagship model-tier arm (gpt-5.4 and gemini-3.1-pro-preview on the frozen 12-specification grid, 1,620 elicitations) finds no evidence that the instability attenuates on larger models. The artifact also derives and validates a deployable specification-instability score (ROC-AUC 0.94 on held-out specifications; 0.70 on the real-issuer arm), with a self-consistency curve and a per-issuer cost model. The reproducible run is offline, deterministic and free; live model capture (vendor API keys) is intentionally excluded.
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.
P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al.· IEEE International Conferenc...· 110 citations· ⚡7
The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· International Conference on...· 84 citations· ⚡6