Sep 2026· Zenodo (CERN European Organization for Nuclear Research)
Advanced Neural Network Applications
Abstract
Hardware-aware quantization frameworks such as HAQ use reinforcement learning (RL)to search per-layer bit-width policies, but this search itself requires hundreds to thousands ofpolicy evaluations with an accelerator (and, in HAQ’s case, per-episode fine-tuning) in theloop — a resource requirement that is at odds with the stated goal of serving teams that lackpowerful accelerators in the first place. We propose replacing the RL search with an LLM-asoptimizer loop, in which a small local language model proposes the next per-layer bit-widthconfiguration directly from a running text log of prior trials, a one-time hardware profile, and aper-layer sensitivity ranking, requiring no training of any kind. We implement this as an endto-end, CI/CD-triggerable pipeline (HAQ-Agent-Lite) that packages the resulting quantizedmodel behind an OpenAI-compatible local inference server, letting a team avoid cloud APIrate limits without owning a powerful GPU. On TinyLlama-1.1B-Chat, run on a laptop CPU,our central empirical finding is not that the LLM optimizer beats uniform quantization inquality — the measured difference (mean proxy score 8.982 ± 0.402 over 4 LLM runs vs. 9.259for uniform 4-bit) is smaller than the LLM’s own run-to-run standard deviation and is notstatistically significant. Rather, we show two things that are supported by the data: (1) ata fixed, small search budget (10 trials), the LLM optimizer is dramatically more reliable thanrandom search over the same space (LLM worst run 9.259 strictly better than random’s bestrun 14.611; Mann-Whitney exact one-sided p = 0.0286), and (2) the LLM systematically avoidscollapsing the top-5 most quantization-sensitive layers to 2-bit (run-level avoidance rate up to10/10 vs. 2–4/10 for random search under the same budget), which we attribute to its use ofthe sensitivity ranking rather than to any implicit training. We further document the searchcost ratio against HAQ’s own released code (600 training episodes + 20 warmup episodes, eachincluding a fine-tuning epoch, vs. our 10 trials with no fine-tuning: a 60× reduction in policyevaluations, though not an apples-to-apples comparison of policy quality or transferability), anegative GPU-offload result (no inference speedup for a 1.1B model on a 4GB Pascal GPU),and several honest approximations imposed by real GGUF quantization kernels, which do notsupport true per-layer mixed precision. We argue the resulting claim — efficient, robust, RL-freeper-layer search at a cost of a handful of LLM calls — is more modest than a quality-superiorityclaim, but is better supported by the evidence and more useful as a deployment strategy forcompute-constrained teams.
Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...
O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al.· Zenodo (CERN European Organi...· 465 citations
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.
Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al.· Journal of Systems and Softw...· 78 citations· ⚡6