Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
Two audit protocols, the comparison of grounding and truth and the swap to an independent evaluator, and RECAP (Readable Encodings via Co-trained Auxiliary Predictors), linear heads trained alongside the target model to keep designated content decodable are contributed.
Hiskias Dingeto
· 0 citations