Background: Pathogenesis-related (PR) proteins, primarily thaumatin-like proteins (TLPs) and chitinases, are the principal cause of protein haze formation in white wines, increasing bentonite requirements and affecting winemaking efficiency. However, evaluating their spatial variability before harvest remains challenging because conventional analytical methods are destructive, labor-intensive, and spatially limited. Methods: This study developed a non-destructive framework to predict the accumulation of PR proteins in Vitis vinifera cv. Chardonnay and Sauvignon Blanc by integrating multitemporal UAV-derived multispectral imagery with stem water potential (Ψstem). Vegetation indices and physiological measurements were acquired throughout berry development over two growing seasons. Elastic Net, XGBoost, and CatBoost models were developed using the 2023 growing season as the calibration dataset through 20 independent jackknife training iterations, with five-fold cross-validation for hyperparameter optimization and an internal 80/20 split used exclusively for early stopping. The models were subsequently externally validated using the independent 2024 growing season. Model interpretation was performed using SHAP to identify influential predictors of model predictions. Results: CatBoost provided the most consistent predictive performance across response variables and was therefore selected for spatial prediction. SHAP analysis revealed cultivar-specific predictor hierarchies, with GNDVI at harvest dominating predictions in Chardonnay, whereas multitemporal NDVI variables were the most influential predictors in Sauvignon Blanc. Stem water potential acquired during Berry Filling II and pre-harvest consistently contributed to model performance in both cultivars, highlighting the importance of late-season physiological conditions. Spatial prediction maps revealed marked intra-vineyard heterogeneity in PR protein accumulation, identifying vineyard sectors with contrasting predicted protein concentrations. Conclusions: Integrating multitemporal UAV multispectral imagery, stem water potential, and explainable machine learning provides an accurate and interpretable framework for predicting the accumulation of pathogenesis-related proteins before harvest. This approach expands the application of remote sensing from conventional assessments of vine vigor to the prediction of biochemical traits directly associated with wine protein stability, supporting targeted sampling, selective harvesting, and more efficient bentonite management in precision viticulture.
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.
Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al.· Journal of Systems and Softw...· 78 citations· ⚡6
This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.
D. Graziotin, Xiaofeng Wang, P. Abrahamsson· SSE@SIGSOFT FSE· 56 citations· ⚡4
This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.
Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson· International Conference on...· 44 citations· ⚡5
It is demonstrated that linker-free PROTACs can outperform traditional designs, marking a paradigm shift in PROTAC development for targeted protein degradation.
Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or sequence constraints.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.