HalluTracer: Pre-Decoding Truthfulness Prediction via Depth-Averaged Probe-Logit
Internal-state probes enable truthfulness prediction before a large language model generates an answer. When detectors change both the layers they read and the rules used to combine them, the source of improved prediction becomes difficult to identify. We separate these choices and find that retaining more layers impro...