Measuring Activation Control in Large Language Models
The Activation Controllability Benchmark is introduced to quantify the extent to which models can modulate their residual stream via natural-language instruction, and suggests that control over the activation space itself could become a confound for monitoring as introspective capabilities increase.