Skip to content

Author

Edwin Marshall III Honeycutt

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Open access Sep 2026

Emulation Diagnostics for Model Self-Reports: A Bridge Between Interpretability and Welfare Assessment

Two adjacent but methodologically separate substrates currently coexist inside Anthropic's published record. The first is a body of mechanistic-interpretability work, most prominently On the Biology of a Large Language Model, that produces attribution graphs over sparse activation features and explicitly cautions that a model's verbal self-explanation can diverge from the underlying mechanism. The second is a welfare-assessment program, most visibly the §5 Claude Opus 4 welfare assessment in the May 2025 system card, that interprets model self-reports about preferences, distress, and affect. The vulnerability this paper takes up does not live inside the welfare assessment itself; it lives in downstream propagation. When the §5.2 descriptive prose is quoted in press, summaries, and secondary commentary stripped of its §5.1 hedge frame, the surface-signature load the §5.1 frame had been carrying is no longer present. This paper proposes four diagnostics: an emulation-versus-state activation contrast over matched prompts; a distribution-dependence ablation re-running the published two-instance free-conversation protocol; a structural-signature comparison of self-reports against other-reports and against empirically calibrated uncertainty; and a propagation diagnostic that compares §5.2-with-hedge-frame against §5.2-hedge-stripped against factual and creative comparators across a six-pattern ghost-pattern taxonomy. The framing is additive to the welfare program rather than adversarial.

Edwin Marshall III Honeycutt · 0 citations
#large language models Open access Sep 2026

The Ghost Pattern Codebook: A Public-Surface Taxonomy of LLM Output Patterns

Large language model outputs fail in ways that the standard evaluation categories, factual error, refusal, hallucination, sycophancy, do not adequately capture. A separate class of failure appears at the level of output form: text that remains fluent, structured, and superficially credible while abandoning the inferential, evidential, or attributional discipline it appears to enforce. The Ghost Pattern Codebook presents a public-surface taxonomy of seven such patterns, Adversarial Loop, Rigorous Wrapper, Formal Dress, Decorative Formalism, Closed Loop, Authority Bleed, and Narrative Pressure, derived from sustained observation of model output in research, scientific-claim filtering, and long-horizon analytical workflows. Each pattern is defined by a behavioral signature: a description of what the failure looks like in the surface text, what evidential operation it imitates, and how it differs from adjacent patterns. The taxonomy is offered as a shared vocabulary for downstream evaluation work, for human reviewers, editors, and AI safety researchers who need to recognize and name these failures without depending on any particular detection tool. The paper documents cross-pattern interactions observed in practice, the decoder-marketing chain, the narrative-pressure attractor, the adversarial-validation cycle, the authority-transfer chain, and the inflation-to-closure lifecycle that recurs in sustained research-mode sessions. The taxonomy is sibling to the epistemic-boundary failure mode formalized in Honeycutt (2026, DOI 10.5281/zenodo.18690241).

Edwin Marshall III Honeycutt · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.