Unsupervised Features Mining via Activation Geometry
The same method is used to select the best training datasets for prompt-injection classifier probes: while similarity between ordinary activations is almost unrelated to downstream performance, RFD-based similarity achieves Top-1 and Top-2 accuracy.