A Stacking-Fusion Feature Selection Framework for Cross-Year and Cross-Cultivar Leaf Hyperspectral Rice Yield Estimation
Single-criterion feature selection for hyperspectral crop yield estimation suffers from methodological bias and limited generalisation. This study proposes a stacking-fusion feature selection framework with a ridge-regression meta-learner (STACKING_FUSION) that transfers the stacked-generalisation concept to the feature-evaluation space, integrating the Pearson correlation coefficient (PCC), grey relational analysis (GRA), and variable importance in projection (VIP) and using the out-of-fold R2 as a meta-supervision signal for adaptive weighting. In field experiments (2024–2025, Gongzhuling, Jilin Province) on rice cultivars Jijing 830 and Jijing 855, leaf hyperspectral reflectance (400–2400 nm) was acquired under controlled indoor measurement conditions at the tillering, jointing, flowering, and milking stages; the study was thus conducted at the leaf scale rather than at the canopy scale of UAV or satellite remote sensing. Second-derivative spectra outperformed original and first-derivative spectra at most stages, and STACKING_FUSION with XGBoost achieved the highest accuracy (R2 = 0.948, RMSE = 0.018 kg m−2, ratio of performance to deviation (RPD) = 4.399). Joint interpretation using SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME), computed on a held-out test subset, identified flowering-stage indices as the leading predictors within the model and suggested a flowering–tillering cross-stage association. With the configuration fixed from the 2024 development dataset, the 2025 analysis showed that the framework was reusable as a locked feature-engineering and algorithmic configuration rather than as a directly portable fitted predictor: under strict zero-shot application it retained high predicted–measured correlations (Pearson r≈ 0.91–0.93) but showed a consistent negative bias together with additional scale and residual error, whereas recalibration on a target-domain calibration subset (70% of the 2025 samples) achieved RPD > 3.0 in both the cross-year and cross-cultivar evaluations. These results indicate that the reusable component is the locked feature-engineering and algorithmic setting rather than the fitted coefficients.