Skip to content
#small language model Open access

OSIRIS v4.5.3: a governed, local-first console for testing whether a small language model learns from conversation

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research) · 2 references
Topic Modeling

Abstract

v4.5.3 is a repair release. The exchange ledger's writer now chains each entry to the ledger file's last entry under a lock, so two consoles can no longer fork the hash chain; the chain check reports forks separately from damage (the author's ledger: 123 entries intact, seven forks from 2026-10-07, nothing rewritten). The Gemini gateway and its tests are restored to v4.5.2 after a merge of a superseded branch left them broken. See RELEASE_NOTES_v4.5.3.md. v4.5.2 corrects how the core is scored. Until this release a layer of the core (the phase-conjugate corrector) ran during training in every block but was skipped whenever gradients were off -- in every held-out score, in early stopping and in generation -- so every held-out figure on record, including those below, describes a different network from the one trained. The layer now runs the same way in both; held-out scores carry the version of the scorer that produced them, older ones no longer count toward the speaking gate, and osiris train --rescore-only replaces them. The release also pre-registers NCLM-ARCH-1, which asks whether the CRSM layers help the core at all against a standard transformer of the same size, and adds claims-register verdicts for the premises of a proposed quantum-LLM roadmap. Details and upgrade steps: RELEASE_NOTES_v4.5.2.md. OSIRIS is a local-first console in which deterministic code mediates every interaction with language models: models propose; code decides, measures and records. Its subject is a small transformer, the core (osiris.nclm), trained online on conversations with a locally hosted mentor model. The core may answer in its own voice only after it passes a held-out speaking gate: 30 held-out exchanges at ≤ 2.0 bits/byte and below a unigram baseline. Until then the mentor answers, labelled as speaking for OSIRIS. Evidence to date: in two pilots, training on conversation lowered the core's loss on held-out replies (mean +0.043 bits/byte on 3 items, 10.5281/zenodo.23075229; +0.067 bits/byte on 20 items, positive on every item, 10.5281/zenodo.23102693). The core nonetheless remained worse than a unigram model of the same replies (5.56 vs 4.68 bits/byte) and has not passed its gate. About 86 % of the variance in the learning signal came from the training run, so a confirmatory test must replicate training runs. The pre-registered confirmatory test (NCLM-1) has not been run. This release claims no successful learning. v4.3.1 is a packaging and repository-hygiene release: modules and scripts the console needs are now packaged (a clean install previously reported 'bench evidence unreadable (ModuleNotFoundError)'), hard-coded home-directory paths are gone, and private conversation transcripts and third-party contact details were removed from the repository. The files of the two previous versions are restricted for that reason. The attached note describes the software, its measurement design and its limitations.

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.