Skip to content

PREreview of "Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring"

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23129468. Summary This paper attacks what it calls the compliance fiction: the industry practice of treating regulatory conformity as a binary verdict declared at deployment time, when agentic and generative systems are dynamic — their behavior evolves with use, context, and model updates. The author introduces governance from metrics, the principle that compliance should be derived as a continuous signal from runtime observability, and presents govllm, an open-source framework implementing governance-driven routing in which model selection is determined by accumulated compliance scores rather than latency or cost alone. Central to the design is a panel of regulatory judges — small language models specialized per criterion (EU AI Act, GDPR, ANSSI, accessibility) — whose inter-judge disagreement is reframed not as noise but as a regulatory uncertainty signal warranting human arbitration. The framework is validated on a ground-truth corpus of 49 annotated prompt/response pairs across five regulatory criteria, evaluated by four small language models (1.7B–7B parameters) running fully on-premise. Agreement rates range from 51.5% (mistral:7b) to 69.1% (phi4-mini), with no single model dominating across all criteria — empirically motivating the Profile-as-jury design. The paper further documents three structural failure modes in small regulatory judges and a judge-specific position bias degrading agreement by up to 25 percentage points across question-order conditions. Strengths The "compliance fiction" framing is the paper's core contribution, and it is exactly right. I audit ML and AI systems for production readiness, and the gap this paper names is one I see constantly: a conformity assessment performed at time t0 is treated as a durable property of the system, while the system's behavior keeps moving. Naming it a fiction — and showing it is structurally incompatible with the EU AI Act's demand for continuous human oversight (art. 14) and lifecycle risk management (art. 9) — gives practitioners language for an argument they currently lose to schedule pressure. This framing alone justifies the paper. On-premise evaluation as regulatory necessity, not design preference. The observation that routing compliance evaluation through external APIs may itself violate the obligations being assessed (GDPR art. 44 on international transfers) is sharp and under-appreciated. Most governance tooling assumes cloud evaluation is available; this paper correctly identifies that for regulated sectors it often is not, and designs around the harder constraint. That is the kind of architectural honesty production work requires. Inter-judge disagreement reframed as signal rather than noise. The existing literature (PoLL, cascaded evaluation) treats variance across judges as something to suppress through aggregation. This paper's move — disagreement among specialized judges marks the compliance grey zone and should trigger human arbitration — is genuinely novel and operationally useful. Finding 5 (bimodal disagreement on adversarial prompts identifying the grey zone more reliably than hard prompts) is the most interesting empirical result in the paper, and it falls directly out of this reframing. The limitations discussion is admirably honest. Section 6.3 quantifies its own weakness with unusual candor: ±30pp confidence intervals at the case level, single-annotator ground truth, 3 of 24 orderings tested, lifecycle drift detection not validated on production data. A paper that tells you exactly where its numbers are soft earns trust for the numbers it stands behind. The open-source release (framework + 49-case corpus) makes the honesty checkable, which is the right combination. Checklist-based validity anchored to jurisprudential sources. Binary checklists tied to CNIL, ANSSI, and AI Act provisions are a credible attempt at making "validity" (does the judge measure compliance?) separable from "reliability" (does the judge agree with itself?). The null self-preference finding — plausibly explained by checklist framing suppressing fluency-preference mechanisms — is a worthwhile, if preliminary, contribution to the judge-bias literature. Major comments/issues 1. The corpus is too small for the comparative claims built on it. 49 cases (10 or fewer per criterion) with the author's own ±30pp case-level confidence intervals means the headline ranking — phi4-mini 69.1% vs. mistral:7b 51.5% — sits inside overlapping uncertainty. The paper is candid about this in section 6.3, but the main text still presents judge rankings, criterion difficulty orderings (Tables 7–8), and the Profile-as-jury optimal assignment as findings rather than hypotheses. Either the corpus needs substantial expansion before these comparisons are made, or every comparative claim needs explicit hedging that survives into the abstract. As written, a reader who skips section 6.3 will come away with rankings the data cannot support. 2. Single-annotator ground truth is the deeper validity problem. The author constructed and annotated all 49 cases alone; inter-annotator agreement (Cohen's kappa) with domain experts is unmeasured. For criteria like human_oversight and non_manipulation — where the paper itself notes the compliant/violating boundary is "inherently contextual" (section 7.4) — one researcher's mapping from legal obligation to binary question is not ground truth, it is one interpretation. The checklist is anchored to regulatory texts, but anchoring is not validation. Expert annotation with reported kappa is, as the author notes, "a necessary step toward a publishable reference benchmark" — I would go further: it is a necessary step before the agreement rates (as opposed to the agreement methodology) can be taken seriously. 3. Position bias is measured on 3 of 24 possible orderings. The "conditional robustness" narrative — phi4-mini immune to reversal but collapsing under permutation — rests on a single permuted ordering (q2→q4→q1→q3). With the full permutation space unexplored, we cannot know whether the observed pattern reflects a general property of the judge or an idiosyncrasy of that one permutation interacting with those specific cases. The finding is intriguing but should be labeled exploratory; the three-ordering design is a pilot, not a characterization. 4. Two of the six claimed contributions lack empirical validation. The governed qualification lifecycle (test → human gate → production → quarantine) and trajectory-based routing are described as implemented and operational but "not empirically validated against production data in this study." A framework paper can certainly propose unvalidated components, but they should then be framed as architectural proposals, not contributions on equal footing with the validated judge panel. The conclusion's claim list should distinguish validated from proposed. 5. The compliance gate needs an error-rate analysis. The per-use-case minimum score threshold automatically excludes underperforming models from routing — "without human intervention on every routing decision." Given judge agreement rates of 50–80%, the gate will produce false exclusions (compliant models routed away, with real latency/cost/availability consequences) at some rate the paper never quantifies. A policy-as-code mechanism that acts on a statistical signal needs its own false-positive analysis. What is the gate's precision, and who reviews its exclusions? 6. No cost or latency analysis of panel judging at production scale. The architecture runs a panel of judges over production interactions to produce continuous compliance scores. Four small language models per interaction, on-premise, is not free — and the paper's own routing proposal would multiply this across use cases. The evaluation literature the paper cites (PoLL) explicitly trades accuracy against a 7x cost reduction; this paper makes no analogous accounting. For a framework whose selling point is operational governance, the operational cost of the governance itself is first-order information. Even rough per-interaction inference costs on the reported hardware would help. 7. The "parameter count is a poor proxy" claim is n=4. Pearson r=-0.39 across four models, with the author's own "indicative only; no statistical significance is implied" caveat, does not support a general claim about scaling and governance. It is consistent with the judge-side finding (phi4-mini beating mistral:7b), and the direction is plausible, but the paper should present this as a suggestion motivating larger-panel studies — which, to be fair, section 7.3 partly does — rather than as an established result. Minor comments/issues - The checklist authorship problem (section 7.4) deserves a concrete next step: the paper could release the annotation guidelines alongside the corpus so independent experts can replicate or contest the mappings, turning a limitation into a community process. - The Incoherence-B false-positive inflation from negation constructions is disclosed honestly; the 29-marker filter reducing phi4-mini's transparency false-positive rate from 21.2% to 13.8% is good practice. But it also suggests the incoherence metric is measuring prompt-phrasing artifacts as much as judge incoherence — worth a sentence on construct validity. - The French-language cases (n=4 of 49) are insufficient to support any bilingual claim; the paper is careful not to overclaim here, but the "intentionally bilingual" framing in section 6.3 slightly oversells n=4. - mistral:7b's order-insensitivity at 51.5% global agreement is characterized as "structural rather than positional" — elegant, but at that agreement level it may simply be a floor effect. A chance-agreement baseline per criterion would clarify whether mistral is judging badly or barely judging at all. - The

View source

Similar papers

#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Conference Sep 2010

Exploring the Sources of Waste in Kanban Software Development Projects

The application of agile software methods and more recently the integration of Lean practices contribute to the trend of continuous improvement in the software industry. One such area warranting proper empirical evidence is a project’s operational efficiency when using the Kanban method. This short paper takes a new an...

Marko Ikonen, Petri Kettunen, Nilay V. Oza et al. · 67 citations · ⚡9

Related blog posts

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.