The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Convergent Category Geometry in Small Language Models
Computer Science
TL;DR
This work extends the framework that demonstrated that truth representations in large language models are universal across statement polarity but reside within a multidimensional subspace along three questions: how the dimensionality of the subspace depends on the model’s knowledge, which architectural component builds the truth direction, and what the direction is a mixture of.