Skip to content
#federated learning Open access

Machine learning-based risk prediction models for type 2 diabetes in primary care: a scoping review

Sep 2026 · Frontiers in Public Health · 23 references
Machine Learning in Healthcare

Abstract

Background Type 2 diabetes mellitus (T2DM) is a major global public health challenge, with many individuals remaining undiagnosed until complications develop. Machine learning (ML)-based risk prediction models have the potential to support early identification of individuals at increased risk using primary care data. However, the characteristics and applicability of these models within primary care settings have not been comprehensively mapped. Objective To systematically map the available evidence on machine learning (ML)-based models for risk prediction, early detection, and case-finding of type 2 diabetes in primary care, and to summarize their characteristics, including predictors, modeling approaches, validation strategies, and model performance. Methods A scoping review was conducted following the Joanna Briggs Institute methodology and reported according to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR). PubMed/MEDLINE, Scopus, and Ovid MEDLINE were searched for English-language studies published between January 2011 and December 2025. Primary research describing ML-based risk prediction models developed, validated, evaluated, or intended for implementation in primary care was eligible. Data were extracted using a structured charting form and synthesized descriptively. Results The search identified 186 records, of which four studies met the inclusion criteria. The included studies were conducted in Sweden, Canada, Saudi Arabia, and Hong Kong between 2024 and 2025. Three studies focused on model development and internal validation, while one externally validated previously developed non-laboratory prediction models. A range of ML approaches was identified, including stochastic gradient boosting, federated learning, multilayer perceptron, random forest, support vector classification, naïve Bayes, and decision tree algorithms, with logistic regression commonly used as a comparator. Models primarily utilized routinely collected demographic, anthropometric, lifestyle, and electronic health record-derived variables. Most studies reported moderate-to-good predictive performance; however, evidence regarding external validation, calibration, and prospective implementation within routine primary care remained limited. Conclusion Evidence on ML-based risk prediction models for T2DM applicable to primary care remains limited despite growing interest in AI for diabetes prediction. Existing models demonstrate promising predictive performance using routinely available clinical information, but greater emphasis is needed on external validation, calibration, prospective implementation, and evaluation across diverse primary care populations before widespread clinical adoption. Review registration Open Science Framework https://osf.io/mbfrz .

View source

Similar papers

#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#machine learning Review Open access Jun 2014

Why Early-Stage Software Startups Fail: A Behavioral Framework

This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.

Carmine Giardino, Xiaofeng Wang, P. Abrahamsson · 175 citations · ⚡19
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#machine learning Review Open access May 2016

Key Challenges in Software Startups Across Life Cycle Stages

It is found that what perceived as biggest challenges by software startups do vary across different life cycle stages, even though its significance decreases when the learning focuses of the startups move from problem to solution and their products mature.

Xiaofeng Wang, Henry Edison, Sohaib Shahid Bajwa et al. · 62 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.