Sleight of Word Benchmark: Can Language Models Notice If Their Own Output Was Tampered With?
A simple benchmark is built in which a single word is consistently substituted with another in the generation process, and two distinct axes are measured: metrics that relate to the model's surprise, as well as an evaluation of the textual reaction for 19 different open-weight language models.
A. Cetoli
· 0 citations