ChatGPT and Generative AI in Physical Education: Applications, Benefits, Limitations, and Risks for Teachers, Students, Coaches, and Professionals
TL;DR
The article proposes the Human-Supervised Generative AI Framework for Physical Education and Sport, which classifies tasks by consequence and requires source grounding, data minimization, professional review, disclosure, and outcome monitoring.
Abstract
Generative artificial intelligence and large language models can draft lesson plans, explain rules, generate assessment materials, support reflective dialogue, summarize training information, and translate specialist language. Their use in physical education, however, differs from use in text-dominant subjects because learning is embodied, safety-sensitive, socially situated, and dependent on observation of movement and context. This structured integrative narrative review synthesizes direct evidence from physical education, sport coaching, exercise science, and sports medicine together with adjacent evidence from education and AI governance. The review distinguishes four user groups: teachers, students, coaches, and exercise or health professionals. Direct studies indicate promising results for structured feedback, rule learning, content knowledge, tactical instruction, lesson planning, and coach reflection, but the evidence base remains small, heterogeneous, and concentrated in short-term or single-context studies. Evaluations of exercise advice show that outputs can be accurate on many general points while remaining incomplete, inconsistent, poorly individualized, or unsafe when medical screening and contraindications are omitted. The article proposes the Human-Supervised Generative AI Framework for Physical Education and Sport, which classifies tasks by consequence and requires source grounding, data minimization, professional review, disclosure, and outcome monitoring. LLMs are most defensible as drafting, explanation, and reflection tools. They should not independently diagnose injury, prescribe rehabilitation, determine return to play, set high-risk training loads, or make consequential student evaluations. The central recommendation is to evaluate AI use by verified educational or professional outcomes rather than fluency, novelty, or time saved.