Preprint
Aug 2026
Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation
This case shows how evaluation-design validity can be checked structurally before model inference and why base correctness does not determine intervention-response fidelity.
Jun Luo, Ning Huang, Ziqi Sha et al.
· 0 citations