Illusions of Accuracy: Widespread Data Leakage Distorts Machine Learning Results in Biochar Research
Abstract
Machine learning (ML) is increasingly used in biochar research to predict yield, material properties, and adsorption performance. Yet, this review reveals that much of the reported success in ML‐driven biochar studies may be overstated due to data leakage, the inadvertent transfer of information between training and testing datasets. We systematically examined 64 of the most‐cited studies published between 2015 and 2025 and found that all studies exhibited potential leakage, commonly arising from hierarchical or grouped data structures, multi‐source compilations, or temporal and concentration dependencies. Nearly all studies (98%) employed random point‐level data splits for training and testing, disregarding dependencies among samples, while only one adopted a group‐based split approach. Re‐analysis of five representative datasets showed that random point‐level splitting can inflate predictive accuracy by up to 86%, often yielding test‐set R2 (coefficient of determination) values exceeding 0.90 even when model generalisation fails. Leakage risks were particularly severe in adsorption studies and in those using very small datasets (< 35 samples), where overfitting further distorted model performance. To ensure methodological rigour and reliable generalisation, biochar ML research must adopt grouped or study‐level data‐splitting strategies, maintain strict separation of preprocessing steps, and incorporate hierarchical or domain‐aware validation frameworks. Future progress requires automated leakage‐detection tools, benchmark datasets with explicit metadata on dependencies, and hybrid models that integrate physical constraints. Addressing data leakage systematically is essential to restore credibility, enhance reproducibility, and realise the full potential of ML as a reliable tool for environmental research.