Skip to content
#small language model Open access

HAQ-Agent-Lite

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)
Advanced Neural Network Applications

Abstract

Hardware-aware quantization frameworks such as HAQ use reinforcement learning (RL)to search per-layer bit-width policies, but this search itself requires hundreds to thousands ofpolicy evaluations with an accelerator (and, in HAQ’s case, per-episode fine-tuning) in theloop — a resource requirement that is at odds with the stated goal of serving teams that lackpowerful accelerators in the first place. We propose replacing the RL search with an LLM-asoptimizer loop, in which a small local language model proposes the next per-layer bit-widthconfiguration directly from a running text log of prior trials, a one-time hardware profile, and aper-layer sensitivity ranking, requiring no training of any kind. We implement this as an endto-end, CI/CD-triggerable pipeline (HAQ-Agent-Lite) that packages the resulting quantizedmodel behind an OpenAI-compatible local inference server, letting a team avoid cloud APIrate limits without owning a powerful GPU. On TinyLlama-1.1B-Chat, run on a laptop CPU,our central empirical finding is not that the LLM optimizer beats uniform quantization inquality — the measured difference (mean proxy score 8.982 ± 0.402 over 4 LLM runs vs. 9.259for uniform 4-bit) is smaller than the LLM’s own run-to-run standard deviation and is notstatistically significant. Rather, we show two things that are supported by the data: (1) ata fixed, small search budget (10 trials), the LLM optimizer is dramatically more reliable thanrandom search over the same space (LLM worst run 9.259 strictly better than random’s bestrun 14.611; Mann-Whitney exact one-sided p = 0.0286), and (2) the LLM systematically avoidscollapsing the top-5 most quantization-sensitive layers to 2-bit (run-level avoidance rate up to10/10 vs. 2–4/10 for random search under the same budget), which we attribute to its use ofthe sensitivity ranking rather than to any implicit training. We further document the searchcost ratio against HAQ’s own released code (600 training episodes + 20 warmup episodes, eachincluding a fine-tuning epoch, vs. our 10 trials with no fine-tuning: a 60× reduction in policyevaluations, though not an apples-to-apples comparison of policy quality or transferability), anegative GPU-offload result (no inference speedup for a 1.1B model on a 4GB Pascal GPU),and several honest approximations imposed by real GGUF quantization kernels, which do notsupport true per-layer mixed precision. We argue the resulting claim — efficient, robust, RL-freeper-layer search at a cost of a handful of LLM calls — is more modest than a quality-superiorityclaim, but is better supported by the evidence and more useful as a deployment strategy forcompute-constrained teams.

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.