Skip to content
Preprint

Persistent Priors, Preserved Targets: A Stroop-Style Paradigm for Lexical Override

Han-yu Wang
May 2026 · 0 citations · 43 references
Computer Science

Abstract

Local definitions can assign a familiar word a temporary meaning while its usual associations remain useful elsewhere. We measure interference from those associations with a matched Stroop-style paradigm. A conflict prompt defines doctor as forest and compares forest with the familiar associate hospital. A neutral control replaces doctor with a semantically weak word in both the definition and query while keeping forest and hospital fixed. All 11 model-level means are positive. Aggregate means are also positive for all four conflict families and prompt formats. When no redefinition is present, a stronger preference for the familiar distractor predicts more interference in arbitrary-semantic, polysemy/entity, and domain-definition remappings, while the antonym slope is null. Separately, we patch neutral-control activations into antonym prompts in five 1B-2B models. Patching the defined word, the target word in the definition, and the later query word together restores almost all of the target-minus-distractor margin lost in conflict (normalized recovery R in [0.92,1.06]). Replacing only that target-word activation with a donor from another item reduces recovery in every tested case. Those donor patches also lower the distractor logit, while the contextual target falls much more than under the same-item patch that restores the margin.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.