TD-GAE: A multimodal large language model framework for adaptive and explainable guided art learning
TL;DR
Findings indicate that TD-GAE can support creative learning as an adaptive, explainable, and pedagogically grounded AI tutoring framework.
Abstract
Multimodal large language models offer strong potential for adaptive and explainable support in creative learning, but many existing systems mainly focus on content generation rather than structured guidance, progress tracking, and interpretable feedback. This paper presents TD-GAE, a Multimodal Large Language Model framework for adaptive and explainable guided art learning. The proposed method combines staged guidance using a directed acyclic graph, learner-state estimation, retrieval-augmented pedagogical context, structured feedback actions, confidence-gated critique, and safety controls to preserve learner agency and originality. Experiments were conducted on six public datasets covering sketch recognition, sketch-photo retrieval, aesthetic assessment, and fine-art style analysis, with comparisons against prompt-only, retrieval-based, planner-only, and agentic baselines. Results show that TD-GAE improves sketch interpretation by 10.2 percentage points on Quick, Draw! and 8.7 points on TU-Berlin over the prompt-only baseline. It also achieves a 22.2% relative mAP improvement in xemplar retrieval over CLIP, a Spearman correlation of 0.587 for aesthetic assessment, and a 40.7% relative improvement in rubric-based learning gains over the next-best baseline. Human evaluation further reports strong ratings for helpfulness, clarity, appropriateness, and perceived usefulness. These findings indicate that TD-GAE can support creative learning as an adaptive, explainable, and pedagogically grounded AI tutoring framework.