Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs
A controlled protocol for evaluating answer stability is introduced: after a model answers a multiple-choice question correctly, it is challenged with a coherent argument for an incorrect option and measured whether the model flips, finding that self-attribution consistently increases flip rates and pooling wrong-answer arguments across models yields stronger adversarial challenges.