Skip to content

Generative AI-Powered Pedagogical Agents in Immersive Environments for Social Sciences and Humanities Education: A Scoping Review

Sep 2026 · Open Science Framework

Abstract

Title: Generative AI-Powered Pedagogical Agents in Immersive Environments for Social Sciences and Humanities Education: A Scoping Review Purpose This research project investigates how generative artificial intelligence (GenAI)-powered pedagogical agents and virtual instructors are being used within immersive virtual reality (VR) and augmented reality (AR) environments, specifically in the context of social sciences and humanities education. While research on generative-AI agents in immersive learning environments has grown rapidly since 2023, no existing systematic or umbrella review has yet mapped this specific intersection — most prior reviews either predate the generative-AI/large-language-model (LLM) era or focus predominantly on STEM, medical, and engineering education. Given that social sciences and humanities education (e.g., history, geography, civics) involves distinctive pedagogical and ethical demands — such as historical empathy, multi-perspective reasoning, and open-ended interpretive dialogue — that differ meaningfully from technical or procedural training domains, the project aims to determine what is currently known about this intersection, how mature the field is methodologically, and where meaningful gaps remain. To address this aim, the study was designed as a scoping review, following Arksey and O'Malley's (2005) five-stage methodological framework and reported according to the PRISMA-ScR (PRISMA Extension for Scoping Reviews) guideline. A scoping-review design was chosen deliberately over a systematic review or meta-analysis because the objective is to map the breadth, characteristics, and trends of an emerging body of literature — rather than to statistically synthesize effect sizes or assess a narrow effectiveness question — which is appropriate given how new and heterogeneous this specific research area still is. Research Questions The project is guided by four research questions: RQ1: What are the design features (embodiment, mode of interaction, role assumed) of GenAI-based pedagogical agents/virtual instructors used in immersive VR/AR environments in social and humanities education? RQ2: What learning outcomes have been reported in studies on these agents, and in which direction do the findings trend? RQ3: At which educational levels and in which social sciences/humanities subfields have these studies been conducted? RQ4: What are the methodological trends and limitations in the field, and what directions are recommended for future research? Methodology A systematic search was conducted across Scopus, Web of Science, and ERIC (August 2026), combining terms related to pedagogical agents/virtual instructors, generative AI/LLMs, immersive VR/AR/XR technologies, and education. The search was restricted to English-language, peer-reviewed journal articles published between 2023 and 2026 — a window chosen to capture the generative-AI/LLM era specifically. Of 105 records initially identified, a multi-stage screening and eligibility process (title/abstract screening, full-text assessment, and data-charting verification) resulted in 9 studies meeting all inclusion criteria. Data extracted from each study included agent design characteristics, technology used, research design, educational level and subject area, reported learning outcomes, and author-stated limitations and future-research recommendations. Findings were synthesized narratively (rather than statistically) around the four research questions and subsequently interpreted through the theoretical lenses of Presence Theory and Embodied Cognition. Expected/Actual Outcomes The review's findings indicate that the included agents are predominantly designed as embodied 3D characters built on GPT-family models, most often assuming peer or mentor roles within VR environments. Reported effects on learning outcomes (motivation, engagement, partner perception, and, in some cases, academic performance) trend positive overall, though effect sizes vary considerably across studies and are notably smaller in the few studies employing control-group comparisons than in single-group, pre-/post-test designs. A key substantive finding is that the existing literature is concentrated almost entirely at the higher-education level and clusters around language education and AI ethics/literacy — it has not yet reached classic social-studies subfields such as history, geography, or civics education, despite the conceptual gap the study set out to address. Interpreted through Presence Theory and Embodied Cognition, the findings further suggest that an agent's educational impact depends less on its technical sophistication (e.g., visual realism) than on whether an appropriate balance between presence and embodiment has been achieved relative to the nature of the learning task — an "embodiment paradox" identified across several included studies. The project's broader contribution is threefold: (1) it provides the field's first dedicated mapping of the generative-AI/LLM generation of pedagogical agents within the social sciences/humanities education context, filling a gap left by earlier, pre-generative-AI-era reviews; (2) it offers a theoretically grounded interpretive lens (Presence Theory/Embodied Cognition) for understanding why and how these agents affect learning, rather than only cataloguing whether they do; and (3) it identifies concrete directions for future research — including extending investigation to K-12 contexts, directly targeting classic social-studies content, adopting more rigorous control-group designs, and incorporating physiological/multimodal measures alongside self-report data. The review also transparently documents its own methodological limitations (a single-researcher screening stage, no prior protocol registration, no formal quality/risk-of-bias appraisal, and a modest final sample of nine studies), consistent with the exploratory nature of scoping reviews and intended to guide readers in appropriately weighing the strength of the evidence presented.

View source

Similar papers

#artificial intelligence Open access May 2023

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations of such models.

Xiaotian Zhang, Chun-yan Li, Yi Zong et al. · 216 citations · ⚡17
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#artificial intelligence Open access Jul 2024

Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval

This work investigates the possibilities of using LLMs in a resume screening setting via a document retrieval framework that simulates job candidate selection and finds that the MTEs are biased, significantly favoring White-associated names in 85% of cases and female-associated names in only 11.1% of cases.

Kyra Wilson, Aylin Caliskan · 131 citations · ⚡8
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8

Related blog posts

MIT News · Artificial Intelligence Sep 14, 2026

New method enables AI for safety-critical situations

The “HardFlow” algorithm could help generative AI models produce high-quality outputs that obey strict requirements when “pretty close” doesn’t cut it.

GPT-Lab Sep 10, 2026

Responsible AI Must Consider Its Afterlife

AI may appear weightless, but every model depends on physical infrastructure. To understand responsible AI, we need to look beyond algorithms and consider the entire lifecycle of the hardware behind them. The post Responsible AI Must Consider Its Afterlife appeared first on GPT-Lab.

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.