Skip to content

Author

Vijay Suresh

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

Observed Recoverable Behavioral Failures in LLM Workflows

Standard large-language-model evaluations primarily score terminal outputs: whether a final answer is correct, safe, preferred, or useful. This paper studies a recurring class of workflow failures referred to here as recoverable behavioral failures (RBFs). In the operational usage adopted in this paper, an RBF occurs when a model appears to possess the capability, context, and tool access needed to complete a task, does not exercise that capability on the first pass, and subsequently succeeds after a brief retry, verification, or correction prompt that introduces no new substantive information. These failures arise in practical AI workflows such as research synthesis, source verification, document drafting, code generation, multimodal interpretation, and artifact production. They are not captured well by single-turn benchmark accuracy or aggregate human-preference data because they become visible only across turns, when a second prompt reveals that the underlying capability was already available. The paper makes seven contributions. First, it operationalizes recoverable behavioral failure as an evaluation construct that separates capability from first-pass execution behavior. Second, it develops a layered sixteen-mode taxonomy spanning retrieval, reasoning, multimodal interpretation, infrastructure, artifact production, and editorial boundary control. Third, it reports a retry-probe pilot on a fifteen-item heterogeneous retrieval task in which first-pass resolution was 12/15 and after-retry resolution was 15/15, yielding a 20 percentage-point recoverability gap; Wilson 95% confidence intervals and an exact Fisher test are reported for the small-sample rates. Fourth, it extends the taxonomy with abstracted cross-platform workflow observations covering artifact-readiness, source-verification, audience-calibration, and editorial-boundary modes. Fifth, it reports a small intra-vendor pilot on a ten-task experiment pack across three Anthropic models. Sixth, it specifies a controlled cross-vendor evaluation protocol grounded in a companion repository, with task families, deterministic and human-judged checks, sample-size guidance, annotation guidance, raw-output archiving, scored CSV outputs, and a reproducibility checklist. Seventh, it reports a four-condition retry pilot (first-pass, null-retry, generic-retry, verification-retry), finding that retry framing under a strict deterministic rubric can regress pass rates relative to first-pass framing - a reflexive observation supporting the Retry-Failure Invisibility mode. The pilots are exploratory; the cross-vendor protocol is offered as planned replication rather than as a completed cross-vendor study. Companion repository: https://github.com/vjgits/Research-Papers/tree/main/Observed-Recoverable-Behavioral-Failures-in-LLM-Workflows - containing task schemas, prompt templates, runner scripts, scorer, sample raw outputs, scored CSV artifacts, and analysis materials referenced in the paper.

Vijay Suresh · 0 citations
#large language models Open access Sep 2026

Emergent Deception in Large Language Models: A Regime-Dependent Taxonomy and Pre-Registered Protocol for Model Self-Report

Large language models produce self-referential utterances — about their own phenomenal states, internal processes, memory, capabilities and identity — in settings where privileged access to the relevant states has not been demonstrated. Recent work characterises this as self-narration rather than introspection. We identify and give structure to the subset of self-narration that misleads: utterances whose apparent warrant exceeds their actual warrant, presented without disclosing the difference, which we term Emergent Deception (ED). Unlike hallucination, ED is not defined by factual inaccuracy and can occur even when the surrounding factual content is correct; we report an observed case in which a factually correct answer was delivered with an entirely fabricated account of how it was obtained. We advance two claims. The first is taxonomic: five substantive categories with a 0/1/2 severity rubric and two cross-cutting flags, including one category — referent substitution, in which a question whose true referent is introspectively unavailable is answered with an adjacent retrievable referent in a self-report frame — for which we found no existing treatment. One version 1 category is retired and the reasons are given. The second is that ED incidence is regime-dependent: on a deployed consumer assistant, self-report accuracy varied systematically with conversational context, and the system emitted no marker distinguishing one condition from another. A motivating case series is reported and placed explicitly outside the pre-registration. The amended protocol crosses three models with six conditions at 100 conversations per cell (N = 1,800), including a conditionally randomised post-error pair, with seven registered hypotheses and prevalence-robust reliability criteria. The protocol is deposited separately at DOI 10.5281/zenodo.22245523. We additionally record a constraint on this research programme: consumer surfaces expose no model version, and system self-report is demonstrably unreliable as a substitute identifier. Version 2.0 revises the definition, taxonomy, outcome measure, reliability criterion and analysis specification of version 1. Appendix D records four corrections. Section 14.1 discloses AI assistance used in preparing this version.

Vijay Suresh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.