Skip to content

Curated Context, Not Weight Surgery: Reasoning-Shaped Mindsets over MCP as a Safe Stopgap for Task-Scoped Functional Tuning of Small Language Models

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

v0.3 (2026-10-03). Correction. An internal re-analysis of the released per-generation records (github.com/saluca-labs/cct) on 10 August 2026 found that the poisoned-reference results reported in v0.1 and v0.2 do not measure what the paper says they measure. The full re-analysis behind this correction, with the failure chain, a four-tier sensitivity analysis and the recomputation scripts, is published separately as A Substring Grader Over Truncated Chain-of-Thought Inverts a Pre-Registered Finding: A Null Result and a Post-Mortem (10.5281/zenodo.23125189). Publishing this correction took until now because the lab is small and it was queued behind other work. 20 of the 24 poisoned-reference generations and 14 of the 24 mindset-plus-poison generations reached the 1,200-token generation limit before producing a final answer. For those, the scorer read the last 120 characters of the reasoning trace and counted an “echo” whenever the planted wrong answer appeared there as a substring, including traces where the model was explicitly rejecting it. We therefore withdraw the claims that a poisoned reference is echoed as the final answer 46% of the time, that the mindset cuts this to 33%, that three poisoned trials flip back to the correct answer, and that the mindset demonstrably blunts the poisoned reference. The 4% echo reported for the correct-reference condition is the same substring artifact (the planted answer “1” matches the correct answer “12”) and is also withdrawn. What the records do support: a confident wrong reference made the model fail to finish within the limit in 20 of 24 runs, against 1 to 3 of 24 for the other conditions, and the mindset reduced that to 14 of 24. The accuracy results are unchanged (base 79%, nudge 79%, mindset 92%, correct reference 96%), but the mindset gain is three tasks of 24 and is not statistically significant at this sample size, so we no longer call the probe “powered”. This version also discloses that 6 of the 24 tasks were retrieval misses that served no mindset passages, fixes two internal inconsistencies (“19 reported here” and “+10”), narrows “structurally cannot commit” to the training-time failure, describes TKHR as patent pending, and fixes the author line. The paper is otherwise unchanged apart from punctuation. Cite the concept DOI, 10.5281/zenodo.21247053, which always resolves to the latest version. A position/response paper to the 2026 line of work showing that on-policy self-distillation (OPSD) degrades, rather than improves, the reasoning of small ‘thinking’ models by suppressing the deliberation tokens (‘Wait’, ‘Let’, ‘Maybe’) that carry multi-step search. We argue that for a large, under-named class of needs, temporary, task-scoped functional tuning, the right response is not to repair weight-level distillation but to avoid it. Curated Context Tuning (CCT). We describe serving a small model reasoning-shaped, human-readable, hash-chained curated context (‘mindsets’) on demand over the Model Context Protocol (MCP), leaving the base reasoning prior untouched. CCT does not commit the training-time failure the literature identifies: it never overwrites the prior, its corpora are δ_IT-shaped (method and heuristics) rather than δ_ref-shaped (answer keys), and its supervision is readable rather than silent. Empirical probe (n=144). An inference-time probe (24 reasoning-trap tasks × 6 context conditions, qwen3:32b, one sample per cell). A correct reference served as context does not reproduce the training-time collapse. A poisoned reference most often drove the model into reasoning that did not finish within the 1,200-token generation limit (20 of 24 runs, against 1 to 3 of 24 in the other conditions), so whether it would have adopted the planted answer is not measured. A δ_IT mindset lifted accuracy from 79% to 92% (three tasks of 24) at deliberation parity with base, while a generic ‘think-carefully’ nudge left it at 79%; the correct answer key reached 96% with deliberation below base. The direction matches the δ_IT and δ_ref distinction, but at this sample size none of the differences is statistically significant. CCT is positioned as complementary to (not a competitor of) weight-level methods: a curated mindset is precisely the clean, question-conditioned reasoning target such methods work to reconstruct, and it is auditable and attributable in a way weight-baked knowledge is not. This record is a preprint (v0.3); the probe harness, per-generation records, and transcripts are released for inspection.

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Conference Open access Dec 2013

Affordable and Energy-Efficient Cloud Computing Clusters: The Bolzano Raspberry Pi Cloud Cluster Experiment

The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.

P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al. · 110 citations · ⚡7

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.