Machine Intuition
Abstract
Can an artificial intelligence system select the right action before it can articulate a valid reason for that action? Addressing this question at the intersection of reasoning, interpretability, agent safety, and latent computation, we develop a non-anthropomorphic definition of machine intuition as pre-explanatory competence: an action-relevant internal state that is sufficiently informative and causally involved to support a correct decision before a faithful natural-language explanation is available. Synthesizing evidence from chain-of-thought prompting, faithfulness interventions, hidden-state probing, activation steering, mechanistic interpretability, and latent-reasoning systems through 8 August 2026, we note that while final answers and tool-use choices can sometimes be decoded from activations before explicit reasoning begins, difficult multi-step problems are often solved during the generated reasoning trace itself - making token-level deliberation computationally consequential rather than merely explanatory. To separate these regimes, we introduce the Action-Explanation Timing (AET) framework and MINT-Eval, a causal evaluation protocol that distinguishes the earliest causally validated action state from the earliest sufficient, faithful, and interventionally supported explanation, while categorizing faithful latent competence, opaque success, rationalized error, and genuinely deliberative reasoning. The resulting answer is qualified: AI can sometimes act correctly before explaining why, but correctness alone does not establish human-like intuition or trustworthy reasoning; because the same timing gap can arise from useful latent computation, learned heuristics, shortcut features, or post-hoc rationalization, explanations must be treated as evidence to test rather than automatic proof of process, requiring high-consequence actions to undergo external verification, authority controls, and causal audits even when the model's first move is right.