Preprint
Jun 2026
Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination
It is shown that a simple, carefully conditioned probe can reliably detect hallucination, and the results suggest that hallucination mitigation is not simply a matter of identifying the right neurons, and point to a deeper separation between what representations reveal and what they allow us to change.
Vijay Vankadaru, Asha Matthews, Tanya Roosta et al.
· 0 citations