A control-plane analysis distinguishing three concepts routinely collapsed in contemporary AI discourse: event verifiability, proxy validity, and governance sufficiency. Taking Reinforcement Learning with Self-Verifiable Rewards and its SpyRL instantiation as its case, it accepts the engineering contribution — an environment that can check a detection outcome against a preassigned latent variable without a human grader — and argues that this relocates rather than resolves the epistemic problem. A reward can be computed deterministically while measuring the wrong construct; a proxy can correlate with human judgment in one setting and fail under distribution shift, collusion, or a change in stakes; and a model can optimize a valid objective while remaining unauthorized to act. It locates precisely where determinism ends inside the method, treats a reward as an evidence packet with provenance, scope and admissibility rather than as a verdict, sets out an eight-layer control architecture separating target definition from proxy construction, validation, update authorization, deployment qualification, runtime permission and audit, and proposes a validation protocol independent of the training loop. Its central claim is that training success does not confer runtime authority. self-verifiable rewards; RLSVR; SpyRL; reinforcement learning; proxy validity; Goodhart's law; construct validity; evidence provenance; pre-execution control; runtime permission; agentic AI
Lara Stuart-Mueller· Zenodo (CERN European Organi...· 0 citations
A control-plane analysis distinguishing three concepts routinely collapsed in contemporary AI discourse: event verifiability, proxy validity, and governance sufficiency. Taking Reinforcement Learning with Self-Verifiable Rewards and its SpyRL instantiation as its case, it accepts the engineering contribution — an environment that can check a detection outcome against a preassigned latent variable without a human grader — and argues that this relocates rather than resolves the epistemic problem. A reward can be computed deterministically while measuring the wrong construct; a proxy can correlate with human judgment in one setting and fail under distribution shift, collusion, or a change in stakes; and a model can optimize a valid objective while remaining unauthorized to act. It locates precisely where determinism ends inside the method, treats a reward as an evidence packet with provenance, scope and admissibility rather than as a verdict, sets out an eight-layer control architecture separating target definition from proxy construction, validation, update authorization, deployment qualification, runtime permission and audit, and proposes a validation protocol independent of the training loop. Its central claim is that training success does not confer runtime authority. self-verifiable rewards; RLSVR; SpyRL; reinforcement learning; proxy validity; Goodhart's law; construct validity; evidence provenance; pre-execution control; runtime permission; agentic AI
Lara Stuart-Mueller· Zenodo (CERN European Organi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.