Oct 2026· Discover Artificial Intelligence· 53 references
Intelligent Tutoring Systems and Adaptive Learning
Abstract
Abstract Physics education research has developed a range of student-centered approaches that emphasize conceptual understanding and problem solving, creating interest in whether large language models (LLMs) can provide complementary, on-demand tutoring support. Locally deployable models are particularly relevant for institutions seeking greater control over data handling and system operation, but their capability relative to commercial models and their reliability in authentic learning interactions remain insufficiently characterized. We introduce mlphys101, a multilingual multiple-choice benchmark of introductory physics questions, and evaluate Gemma 3 27B on its German variant in three quantization configurations. The locally deployed model achieves high benchmark accuracy, although performance is lower and more variable on conceptual and multi-step reasoning tasks than that of a commercial cloud baseline. We subsequently deploy the same model as a tutoring chatbot in a classroom field experiment with first-semester engineering students. Survey responses indicate that explanations were generally comprehensible, but the practical usefulness of the system was limited by frequent prompt reformulation, occasional incorrect or incomplete responses, difficulties in using information from the ongoing interaction, and response latency. These findings demonstrate that strong performance on a controlled physics benchmark does not necessarily translate into reliable tutoring behavior. Locally deployed LLMs may therefore provide useful support for selected learning activities, but their suitability for educational use should be evaluated not only through domain benchmarks but also under realistic, multi-turn interaction conditions.
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.
P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al.· IEEE International Conferenc...· 110 citations· ⚡7
The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· International Conference on...· 84 citations· ⚡6
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 8, 2026