Aug 2026· International Journal of Data Science and Analysis· Vol 22· 0 citations· 139 references
TL;DR
Eleven practical tips for reducing unintentional overfitting in supervised biomedical machine learning studies are presented and stress principled data splitting, domain-informed preprocessing, controlled model complexity, systematic tuning, comprehensive performance evaluation, and robustness analysis.
Abstract
Overfitting is the excessive adaptation of a machine learning model to its training data and remains a persistent challenge in biomedical informatics. In supervised learning, models may capture patterns “too well,” failing to generalize to unseen data and yielding overly optimistic performance estimates. The problem is especially acute in biomedical settings, where datasets are high-dimensional, heterogeneous, and often of few samples. Numerous strategies have been proposed to mitigate overfitting in bioinformatics and health informatics. However, even established techniques can produce misleading results if misapplied, for example through data leakage, excessive hyperparameter tuning, inappropriate preprocessing, or inadequate validation. To address these pitfalls, we present eleven practical tips for reducing unintentional overfitting in supervised biomedical machine learning studies. The recommendations stress principled data splitting, domain-informed preprocessing, controlled model complexity, systematic tuning, comprehensive performance evaluation, and robustness analysis. Rather than offering an exhaustive treatment, we provide an accessible, practice-oriented guide to support more reliable and reproducible machine learning research. Although developed for biomedical informatics, these quick tips are broadly applicable across disciplines using supervised machine learning.
Ten tips for successfully and sustainably implementing Federated learning for Biomedical applications, ensuring both ethical data governance and improved model performance in sensitive domains are outlined.
Kyle Ellrott, V. Malladi, J. Bélisle-Pipon et al.· PLoS Computational Biology· 0 citations
This work proposes post-pretrained lasso selective inference (PPL-SI), a novel selective inference method designed to provide statistically valid p values for the pretrained lasso that reliably controls false discoveries and significantly improves the detection of biologically relevant features compared to traditional approaches.
Cao Huyen My, Nguyen Vu Khai Tam, Vo Nguyen Le Duy· Statistics and computing· 0 citations
Data serves as the foundation of contemporary artificial intelligence (AI) systems, yet ethical and practical constraints often limit the availability and usability of real-world datasets. Synthetic data (SD) has emerged as a valuable solution, enabling the development and training of AI models without compromising privacy standards or ethical guidelines. However, many challenges remain to address, from generating high-quality SD using various approaches to investigating the impacts of data training on machine learning (ML) models. This study examines the impact of balancing real and SD and provides some recommendations that researchers can further utilise to improve the ML model’s training process. Three datasets, Mobile Health (MHealth), High-Energy Physics Mass (HEPMass), and US Company Bankruptcy Prediction (UCBP), were pre-processed to ensure compatibility and used as the basis for SD generation using Conditional Tabular Generative Adversarial Networks (CTGAN) and Tabular Variational Autoencoders (TVAE). The study employed a hybrid data generation approach, splitting the training data into varying proportions of real and SD, with performance evaluated through 5-fold cross-validation on Decision Tree (DT), Gaussian Naive Bayes (GNB), and Linear Support Vector (L-SVM) ML models. The results indicate the optimal balance of real and SD for maximising model performance. Analysis of benchmarking results across three datasets shows that combining 30% real data with 70% CTGAN-generated synthetic data achieves the highest accuracy and overall model performance. In contrast, when using TVAE-generated data, a 20% real and 80% synthetic split is recommended to maintain similar performance. Additional experiments conducted on 10% reduced subsets showed that while the primary trends persisted under limited-data conditions, the optimal real-to-synthetic data ratio became more sensitive to the specific dataset. Statistical significance of the model performance differences was further confirmed using paired t-tests across all evaluated mixing ratios. This research analyses these results and considers the state of the art to recommend how further synthetic data can be useful with machine learning models and which approaches to adopt.
Majid Liaquat, Chris D. Nugent, Ian Cleland et al.· IEEE Access· 0 citations
Background: Diabetes mellitus is a major global health burden, and its early detection is essential for preventing serious complications. Machine learning, and eXtreme Gradient Boosting (XGBoost) in particular, performs strongly on routine clinical data; however, reported results are frequently optimistic because of hold-out evaluation and data leakage, and the resulting models are often difficult to interpret.
Objective: To develop and rigorously evaluate an interpretable XGBoost model for predicting the presence of diabetes from eight routine diagnostic measurements, using an unbiased, leakage-controlled evaluation design.
Methods: An openly available dataset of 1,168 patient records (771 non-diabetic and 397 diabetic; an approximately 66:34 class imbalance) with eight clinical features was analysed. Physiologically implausible zero values were treated as missing and imputed with the median inside a processing pipeline. Model selection and performance estimation were separated using a 5×5 nested cross-validation scheme, with all preprocessing confined to the training folds to prevent data leakage. Performance was assessed on out-of-fold predictions using accuracy, sensitivity, specificity, precision, F1 score, the area under the ROC curve (ROC-AUC), and the Brier score; probability calibration was examined; and the contribution of each feature was quantified with SHapley Additive exPlanations (SHAP).
Results: The model achieved an ROC-AUC of 0.823, an accuracy of 0.769, a specificity of 0.859, a sensitivity of 0.595, a precision of 0.684, an F1 score of 0.636, and a Brier score of 0.160. The near-diagonal calibration curve and low Brier score indicated that the predicted probabilities were well calibrated and could be interpreted as reliable risk estimates. SHAP analysis identified glucose, body mass index, and age as the most influential predictors, in agreement with established clinical risk factors for diabetes.
Conclusions: Under a leakage-controlled, unbiased evaluation, XGBoost provided moderate but trustworthy discrimination together with well-calibrated probabilities for diabetes prediction, while SHAP confirmed clinically plausible predictors. The comparatively conservative performance underscores the importance of nested cross-validation over simple hold-out splits for realistic model assessment. External validation and decision-threshold or class-imbalance strategies represent promising directions for future work.
Z. Kucukakcali, I. Cicek· International Journal of Med...· 0 citations
The proposed GPT2-based table-to-text framework provides a practical and clinically interpretable approach for disease prediction from limited structured healthcare data and demonstrates strong potential for early risk detection, transparent clinical decision support, and reliable deployment in real-world low-resource healthcare environments.
S. Bin Akter, S. Akter, D. Eisenberg et al.· medRxiv· 0 citations
Risk prediction tools are becoming increasingly popular tools to assist in making clinical decisions. The models, however, are typically trained on data from general patient cohorts and may not be representative of and applicable to targeted patient cohorts when used in practice. This study overcame these obstacles by developing and evaluating a clinical risk prediction model using the MIMIC-III clinical dataset and an Artificial Neural Network (ANN). The suggested method accounts for all possible preprocessing steps—including normalization, encoding, managing missing values, and data balancing using SMOTE-ENN—to increase the dependability of the predictions. With an F1-Score (F1) of 96.9%, an accuracy (ACC) of 98.6%), a precision (PRE) of 97%, and a recall (REC) of 96.5%, the ANN model has high predictive capacity and can capture the complicated interaction between the clinical variables. The outcomes demonstrate that the suggested ANN architecture can accurately and consistently assess clinical risk, which in turn allows for the early identification of high-risk patients and aids healthcare providers in making data-informed therapeutic decisions. Because it can provide trustworthy prediction models from complex health data, the proposed method has great potential for clinical real-time application. It can also help develop smarter healthcare decision-support systems and enhance patient monitoring methods.
Shivani Jain· Journal of Artificial Intell...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.