Oct 2026· Zenodo (CERN European Organization for Nuclear Research)· 2 references
Topic Modeling
Abstract
v4.5.3 is a repair release. The exchange ledger's writer now chains each entry to the ledger file's last entry under a lock, so two consoles can no longer fork the hash chain; the chain check reports forks separately from damage (the author's ledger: 123 entries intact, seven forks from 2026-10-07, nothing rewritten). The Gemini gateway and its tests are restored to v4.5.2 after a merge of a superseded branch left them broken. See RELEASE_NOTES_v4.5.3.md. v4.5.2 corrects how the core is scored. Until this release a layer of the core (the phase-conjugate corrector) ran during training in every block but was skipped whenever gradients were off -- in every held-out score, in early stopping and in generation -- so every held-out figure on record, including those below, describes a different network from the one trained. The layer now runs the same way in both; held-out scores carry the version of the scorer that produced them, older ones no longer count toward the speaking gate, and osiris train --rescore-only replaces them. The release also pre-registers NCLM-ARCH-1, which asks whether the CRSM layers help the core at all against a standard transformer of the same size, and adds claims-register verdicts for the premises of a proposed quantum-LLM roadmap. Details and upgrade steps: RELEASE_NOTES_v4.5.2.md. OSIRIS is a local-first console in which deterministic code mediates every interaction with language models: models propose; code decides, measures and records. Its subject is a small transformer, the core (osiris.nclm), trained online on conversations with a locally hosted mentor model. The core may answer in its own voice only after it passes a held-out speaking gate: 30 held-out exchanges at ≤ 2.0 bits/byte and below a unigram baseline. Until then the mentor answers, labelled as speaking for OSIRIS. Evidence to date: in two pilots, training on conversation lowered the core's loss on held-out replies (mean +0.043 bits/byte on 3 items, 10.5281/zenodo.23075229; +0.067 bits/byte on 20 items, positive on every item, 10.5281/zenodo.23102693). The core nonetheless remained worse than a unigram model of the same replies (5.56 vs 4.68 bits/byte) and has not passed its gate. About 86 % of the variance in the learning signal came from the training run, so a confirmatory test must replicate training runs. The pre-registered confirmatory test (NCLM-1) has not been run. This release claims no successful learning. v4.3.1 is a packaging and repository-hygiene release: modules and scripts the console needs are now packaged (a clean install previously reported 'bench evidence unreadable (ModuleNotFoundError)'), hard-coded home-directory paths are gone, and private conversation transcripts and third-party contact details were removed from the repository. The files of the two previous versions are restricted for that reason. The attached note describes the software, its measurement design and its limitations.
Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...
O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al.· Zenodo (CERN European Organi...· 465 citations
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.
Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al.· Journal of Systems and Softw...· 78 citations· ⚡6