Skip to content

HAQ-Agent-Lite

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)
Advanced Neural Network Applications

Abstract

Hardware-aware quantization frameworks such as HAQ use reinforcement learning (RL)to search per-layer bit-width policies, but this search itself requires hundreds to thousands ofpolicy evaluations with an accelerator (and, in HAQ’s case, per-episode fine-tuning) in theloop — a resource requirement that is at odds with the stated goal of serving teams that lackpowerful accelerators in the first place. We propose replacing the RL search with an LLM-asoptimizer loop, in which a small local language model proposes the next per-layer bit-widthconfiguration directly from a running text log of prior trials, a one-time hardware profile, and aper-layer sensitivity ranking, requiring no training of any kind. We implement this as an endto-end, CI/CD-triggerable pipeline (HAQ-Agent-Lite) that packages the resulting quantizedmodel behind an OpenAI-compatible local inference server, letting a team avoid cloud APIrate limits without owning a powerful GPU. On TinyLlama-1.1B-Chat, run on a laptop CPU,our central empirical finding is not that the LLM optimizer beats uniform quantization inquality — the measured difference (mean proxy score 8.982 ± 0.402 over 4 LLM runs vs. 9.259for uniform 4-bit) is smaller than the LLM’s own run-to-run standard deviation and is notstatistically significant. Rather, we show two things that are supported by the data: (1) ata fixed, small search budget (10 trials), the LLM optimizer is dramatically more reliable thanrandom search over the same space (LLM worst run 9.259 strictly better than random’s bestrun 14.611; Mann-Whitney exact one-sided p = 0.0286), and (2) the LLM systematically avoidscollapsing the top-5 most quantization-sensitive layers to 2-bit (run-level avoidance rate up to10/10 vs. 2–4/10 for random search under the same budget), which we attribute to its use ofthe sensitivity ranking rather than to any implicit training. We further document the searchcost ratio against HAQ’s own released code (600 training episodes + 20 warmup episodes, eachincluding a fine-tuning epoch, vs. our 10 trials with no fine-tuning: a 60× reduction in policyevaluations, though not an apples-to-apples comparison of policy quality or transferability), anegative GPU-offload result (no inference speedup for a 1.1B model on a 4GB Pascal GPU),and several honest approximations imposed by real GGUF quantization kernels, which do notsupport true per-layer mixed precision. We argue the resulting claim — efficient, robust, RL-freeper-layer search at a cost of a handful of LLM calls — is more modest than a quality-superiorityclaim, but is better supported by the evidence and more useful as a deployment strategy forcompute-constrained teams.

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#machine learning Review Open access Jun 2014

Why Early-Stage Software Startups Fail: A Behavioral Framework

This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.

Carmine Giardino, Xiaofeng Wang, P. Abrahamsson · 175 citations · ⚡19
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15

Related blog posts

GPT-Lab Aug 28, 2026

We built an AI factory for HVAC control

What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.

Microsoft Research Blog Jul 30, 2026

EvoLib: Turning experience into evolving knowledge

LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib: Turning experience into evolving knowledge appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.