The Mechanism Before the Mind: An Engineering Standard for Causal Claims About Machine Learning Behavior
Machine learning research routinely treats a behavioral description as if it were a demonstrated internal mechanism, calling something "intention" or "preference" and then reasoning as though that label had been independently established rather than just applied. This paper draws that boundary formally: an observed behavior (C → B) and a proposed causal mechanism (C → I → B) are different evidentiary claims, and the second requires independent identification the first does not. It proposes a nine-question audit for testing whether any given psychological construct in the AI literature has actually cleared that bar, or is only naming the behavior it's meant to explain.