Skip to content
Open access

Integrating machine learning, deep learning, and image analysis for seed species classification

Jul 2026 · Applications in Plant Sciences · 0 citations · 25 references

Abstract

The growing demand for wildflower seeds in ecological restoration requires reliable species identification, yet current market products often contain heterogeneous species. As seed identification is labor‐intensive and requires advanced botanical knowledge, we evaluated multiple segmentation and classification approaches to determine which combination performed better for identifying species and quantifying their frequencies under real‐life constraints. We generated seed images from common wildflower species using a flatbed scanner. Artificial intelligence (AI) and non‐AI segmentation pipelines were compared. We evaluated over 20 classifiers on tabulated features and convolutional neural networks (CNNs) on single seed image inputs. Open‐set conditions were simulated using “mock seed mixes” with previously unseen taxa using decision thresholds. With tabulated data, the XGBoost and Multi‐layer Perceptron (MLP) models exceeded F1 > 0.96, while AutoGluon's CNN and ensemble models reached F1 > 0.97. ResNet‐50 achieved an F1 score over 0.99 on single‐seed images obtained with Cellpose segmentation. Random forest achieved the highest accuracy (0.92) with unseen species in open‐set classification. Deep learning segmentation substantially enhanced CNN accuracy, yet tabulated‐feature machine learning models remained competitive and efficient. Threshold‐based rejection enabled robust open‐set classification, with random forest outperforming deeper models in rejecting unseen seeds. No single method was universally optimal; the best strategy depends on whether the task involves closed‐set or open‐set classification and on trade‐offs between accuracy and computational cost.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.