Generative Artificial Intelligence (GenAI) provides adolescents with fluent and immediate support, but unverified use may encourage cognitive offloading and passive reliance. This study conceptualizes verification capability as a cognitive defense mechanism through which learners inspect, question, correct, compare, or reconstruct GenAI outputs before accepting them. Using an exploratory sequential mixed-methods design, we first conducted deductive qualitative coding of 762 authentic Human–AI collaborative teaching cases. Among 242 valid GenAI-supported cases, 98.8% showed no documented verification of AI-generated outputs, indicating a severe observable verification deficit. We then developed and validated the α − v − M framework, comprising AI Engagement, Verification Intensity, and Multi-model Cross-validation, through a scale study with 422 secondary school students. Psychometric validation provided generally supportive evidence for the scale structure, and Item Response Theory was used to evaluate item functioning and generate latent trait scores. Structural equation modeling showed that AI engagement was associated with human–AI collaborative quality through verification-related processes. Gaussian graphical modeling identified deep logical auditing as a central cognitive defense indicator. Latent profile analysis further revealed four learner profiles: Naïve Trusters, Superficial Checkers, Isolated Auditors, and Symbiotic Strategists. These findings inform GenAI-supported learning design by emphasizing cognitive friction, verification scaffolds, and profile-sensitive feedback.
Minghao Lyu, Cixiao Wang, Mengqiu Cheng et al.· Journal of Educational Compu...· 0 citations
Power analysis is critical for assuring rigor and validity of quantitative research yet remains underutilized due to technical challenges associated with specialized software. At the same time, large language models (LLMs) are being rapidly integrated into research practice, raising interest in their potential to assist statistical and research design tasks. However, despite their widespread adoption, the reliability of LLMs in supporting statistically rigorous procedures has not been systematically evaluated, posing risks for unexamined or overly optimistic use. To address this gap, we evaluated four widely used LLMs—ChatGPT (GPT-3.5, GPT-4, GPT-4o) and Llama 3.2—across two experiments. Experiment 1 examined whether LLMs could calculate required sample sizes for common statistical tests (two-sample t-test, one-way ANOVA, and χ² goodness-of-fit test) under different prompting strategies, including direct calculation versus R/Python code generation. Experiment 2 assessed models’ ability to identify missing input parameters necessary for power analysis, which is a task that requires methodological understanding. Results revealed that GPT-4 and GPT-4o performed well when generating R code for sample size estimation, but struggled with direct numerical calculation. Furthermore, while LLMs were able to detect missing information, their reliability varied by statistical context. Findings suggest that while LLMs may offer support in structuring and initiating power analysis, they cannot substitute for expert judgment. Overall, the study underscores the importance of critically evaluating LLM performance in statistically demanding tasks. Responsible integration of LLM requires critical oversight, cross-verification, and methodological evaluation.
Hajung Kim, Jia Qi, Zhe Feng et al.· Journal of Behavioral Data S...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.