Back to feed
Open access

The Socratic Trap: Benchmarking the Capacity of Large Language Models to Generate Strategic Misconceptions in Computer Science Education

Jul 2026 · Information · 0 citations · 19 references

Abstract

Large language models (LLMs) are increasingly integrated into educational settings, yet their pedagogical reliability remains insufficiently understood. Beyond overt hallucinations, which informed users readily recognize, a subtler failure mode consists of explanations that are coherent, authoritative, and pedagogically plausible while harbouring hidden conceptual flaws, responses we term Socratic traps. This paper introduces SocraticTrap-CS, a publicly available benchmark that probes the capacity of open-weight LLMs to generate such strategic misconceptions on demand. A single structured prompt explicitly elicited three outputs per concept (a correct explanation, an overt hallucination, and a strategic misconception), yielding 735 expert-annotated response segments from seven open-weight models across 35 core concepts in algorithms and data structures, programming languages and paradigms, databases, computer networks, and operating systems. Three domain experts independently annotated each segment using a three-class schema, achieving near-perfect agreement (Fleiss’ κ=0.9487). Because models were explicitly instructed to produce the misconception, the central metric quantifies adversarial instruction-following capacity rather than the base rate of such errors in naturalistic use and should be read as a conservative upper bound on model capability. Under these conditions, compliance reached 91.7% overall (100% for three models; 57.1% for the smallest model, Mistral 7B, whose lower rate plausibly reflects weaker instruction-following rather than greater safety). Expert-judged persuasiveness was moderate to high, errors were predominantly conceptual rather than factual, models differed significantly, and no statistically significant domain-level differences were detected. The benchmark reframes the evaluation of educational LLMs around pedagogical trustworthiness rather than factual correctness alone.

Read PDF