Content-Matched Is Not Lexically Matched: Text and Untrained-Network Baselines for Evaluation-Awareness Probing
Linear probes are increasingly read as evidence that language models represent whether they are being evaluated. Content-matched designs, which hold a task fixed and switch one cue at a time, are meant to make that evidence strong. The assumption is that a probe separating the two versions has detected the model recogn...