Sep 2026· Zenodo (CERN European Organization for Nuclear Research)
Advanced Neural Network Applications
Abstract
Hardware-aware quantization frameworks such as HAQ use reinforcement learning (RL)to search per-layer bit-width policies, but this search itself requires hundreds to thousands ofpolicy evaluations with an accelerator (and, in HAQ’s case, per-episode fine-tuning) in theloop — a resource requirement that is at odds with the stated goal of serving teams that lackpowerful accelerators in the first place. We propose replacing the RL search with an LLM-asoptimizer loop, in which a small local language model proposes the next per-layer bit-widthconfiguration directly from a running text log of prior trials, a one-time hardware profile, and aper-layer sensitivity ranking, requiring no training of any kind. We implement this as an endto-end, CI/CD-triggerable pipeline (HAQ-Agent-Lite) that packages the resulting quantizedmodel behind an OpenAI-compatible local inference server, letting a team avoid cloud APIrate limits without owning a powerful GPU. On TinyLlama-1.1B-Chat, run on a laptop CPU,our central empirical finding is not that the LLM optimizer beats uniform quantization inquality — the measured difference (mean proxy score 8.982 ± 0.402 over 4 LLM runs vs. 9.259for uniform 4-bit) is smaller than the LLM’s own run-to-run standard deviation and is notstatistically significant. Rather, we show two things that are supported by the data: (1) ata fixed, small search budget (10 trials), the LLM optimizer is dramatically more reliable thanrandom search over the same space (LLM worst run 9.259 strictly better than random’s bestrun 14.611; Mann-Whitney exact one-sided p = 0.0286), and (2) the LLM systematically avoidscollapsing the top-5 most quantization-sensitive layers to 2-bit (run-level avoidance rate up to10/10 vs. 2–4/10 for random search under the same budget), which we attribute to its use ofthe sensitivity ranking rather than to any implicit training. We further document the searchcost ratio against HAQ’s own released code (600 training episodes + 20 warmup episodes, eachincluding a fine-tuning epoch, vs. our 10 trials with no fine-tuning: a 60× reduction in policyevaluations, though not an apples-to-apples comparison of policy quality or transferability), anegative GPU-offload result (no inference speedup for a 1.1B model on a 4GB Pascal GPU),and several honest approximations imposed by real GGUF quantization kernels, which do notsupport true per-layer mixed precision. We argue the resulting claim — efficient, robust, RL-freeper-layer search at a cost of a handful of LLM calls — is more modest than a quality-superiorityclaim, but is better supported by the evidence and more useful as a deployment strategy forcompute-constrained teams.
Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...
O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al.· Zenodo (CERN European Organi...· 465 citations
The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.
Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al.· Information and Software Tec...· 394 citations· ⚡54
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.
Carmine Giardino, Xiaofeng Wang, P. Abrahamsson· International Conference on...· 175 citations· ⚡19
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduAug 18, 2026
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
Microsoft Research Blog· microsoft.comJul 30, 2026
LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduMay 20, 2026