Aug 2026· International journal of computer information systems and industrial management applications· Vol 18, pp. 816-830· 0 citations
TL;DR
The article concludes that formal validity and practical rationality answer to different criteria of success, that the empirical superiority of statistical prediction over human judgment does not collapse this distinction, and that the operative difference lies in the structure of failure rather than in its frequency.
Abstract
Algorithmic failures are conventionally diagnosed as deviations from a correct specification — bias to be corrected, error to be patched. This article argues that the more consequential and less visible failure mode runs the other way: a system executes its specification exactly, and the specification was inadequate to the situation it governed. The distinction is developed through a formal argument rather than an analogy. Treating the feature set available to a learning system as a -algebra , the article shows that no -measurable function can represent a value-relevant event excluded from , and — more precisely — that no -measurable self-assessment functional can take the adequacy of itself as an argument. Every reliability measure a model reports, from calibration to conformal coverage, is computed inside the representation whose adequacy is in question. Sections 2 and 3 locate this claim against the literatures on bounded rationality, rule-following, and algorithmic fairness, arguing that each stops short of the reflexive point at issue. Sections 4 through 6 develop the argument through decision theory, the Rashomon set, Goodhart-type selection effects, and an impossibility theorem for fairness criteria, and confront the two strongest objections available — that scope-monitoring is already being formalised, and that closure under a fixed representation cannot be a principled limit if human reasoners are themselves physical systems. The article concludes that formal validity and practical rationality answer to different criteria of success, that the empirical superiority of statistical prediction over human judgment does not collapse this distinction, and that the operative difference lies in the structure of failure rather than in its frequency: formal errors are quiet, mechanically correlated, and reproduced at scale, while human error — correlated though it often is by institutional bias, professional norms, and bureaucratic routine — remains comparatively noisy, locally visible, and costly to scale.
In May 2026 an OpenAI model produced a counterexample to the Erd\H{o}s unit distance conjecture. Five mathematicians published a human-verified version the same day, and the result entered the literature within weeks. In August 2026 the same laboratory published ten mathematical and theoretical computer science results, each accompanied by a machine-checkable Lean 4 certificate with no unproved steps. Four weeks later, one remained the subject of an unresolved dispute over whether its formalization meant what it claimed. We argue that this difference is structural. We distinguish three layers of verification: derivational validity, which a kernel checks; representational fidelity, whether the formal statement means the intended question; and epistemic significance. Only the first is mechanizable. Making it effectively free therefore does not eliminate verification work but shifts the burden to layers dependent on scarce expert attention. Measurements of the August corpus illustrate the shift. The kernel-checked proofs total 20.6 MB, while the statements requiring human audit total 55.6 KB, a ratio of 379 to 1. Yet those statements contain 218 bespoke definitions rather than relying on community-vetted ones. The audit surface is therefore small in volume but irreducibly expert. We argue that machine checking produces verification abundance while leaving adjudication scarce. We propose a six-category taxonomy of representational mismatch, a disclosure schema for machine-generated mathematical claims, and implications for software, cryptography, and regulated decision systems.
The paper formulates the Control Responsibility Principle (CRP), which shifts AI ethics from a focus on what systems know to a focus on how their behavior is governed, who holds authority over that governance, and how such authority is to be justified, distributed, and contested.
An interface audit for distributed AI evidence: typed source and target descriptions, and a procedure separating endpoints that never meet from endpoints that meet while warrant fails to cross, and the resulting projectibility audit diagnoses unsupported joins in benchmark-to-use arguments.
The EU AI Act requires providers of high-risk systems to file technical documentation describing how the system reaches its decisions. Mechanistic interpretability is the obvious source of such evidence, and circuit discovery is its most developed instrument. We ask whether that evidence survives the condition under which it would be relied upon: two competent analysts, the same system, the same tool, different defensible settings. We pre-registered a crossed grid of seven analytic axes, every level taken from a published implementation, and mapped each discovered circuit through a deterministic claim map to a structured Annex IV statement. Across 15,840 pre-registered specifications on GPT-2 small and the indirect object identification task, of which 7,561 produced a claim, the derived statement flips across 73.2% of specification pairs (95% CI 0.725 to 0.738) and the modal claim commands 41.1% of the space. The evidence fails a filability criterion at every tolerance a conformity assessment body would plausibly accept. Standardising the single most influential choice, the evaluation metric, leaves the flip rate at 59.4%. Removing circuit size from the claim entirely and holding it fixed leaves 27.1% (95% CI 0.255 to 0.286), still above the pre-registered threshold. The circuits underlying these claims are structurally near-disjoint, median pairwise Jaccard overlap 4%, and functionally uncorrelated at Cohen's kappa 0.015, so the instability is not one mechanism described in different words. We give the filability criterion as a standalone protocol, and we report that one of the seven documented discovery objectives does not execute at all on the library's own canonical task. The study covers one model and one task, and whether the conclusion holds at scale is untested.
This work defines the evidential ceiling of an evaluation as the largest factor by which one result can move belief under a fixed testing budget, derive it in closed form for the benchmark null result, and uses it to locate that boundary exactly.
A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.
V. Rodionov, Shamil Assylbekov· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.