Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity
This work probes models to answer two questions: do their activations encode whether a referent falls inside the knowledge boundary, and do they anticipate the specificity of the referent they are about to generate?