Oct 2026· Journal of Chemical Information and Modeling· 0 citations· 24 references
Abstract
Low-populated protein excited states can play decisive roles in function, ligand recognition, and misfolding, but their structural characterization remains difficult because these conformations are transient and sparsely populated. Here, we present CS-SelectFold, a framework that combines generative conformational sampling with NMR chemical-shift-guided post hoc selection to recover low-populated protein conformations without additional training or fine-tuning of its pretrained components. CS-SelectFold uses SimpleFold to sample diverse candidate structures from sequence, applies structural prescreening to enrich models that escape the dominant ground-state basin, predicts chemical shifts with UCBShift 2.0, and ranks candidates by agreement with experimentally determined excited-state chemical shifts. We benchmarked CS-SelectFold on two proteins with experimentally characterized low-populated states. For the T4 lysozyme L99A cavity mutant, a canonical excited-state system, the method recovered the defining excited-state-like local rearrangement around the engineered cavity, including inward placement of Phe114, consistent with the NMR/CS-Rosetta-derived excited-state model. A retrospective two-reference hotspot analysis using internal-distance RMSD (dRMSD)─the root-mean-square difference between corresponding pairwise distances within the hotspot rather than conventional coordinate RMSD after structural superposition─further showed that chemical-shift-based ranking enriched conformations that moved away from the ground-state basin and toward the experimentally defined excited-state basin. For the A39V/N53P/V55L Fyn SH3 domain, which populates an aggregation-prone hidden folding intermediate, CS-SelectFold recovered the key topological signature of the experimentally characterized intermediate: loss of the C-terminal β-strand present in the native state, as confirmed by both three-dimensional structural comparison and secondary-structure analysis. These results provide proof-of-concept support for CS-SelectFold and suggest that the same sample-and-select principle may extend to broader low-populated hidden conformations, including folding intermediates.
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.
Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al.· Journal of Systems and Softw...· 78 citations· ⚡6
This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.
D. Graziotin, Xiaofeng Wang, P. Abrahamsson· SSE@SIGSOFT FSE· 56 citations· ⚡4
This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.
Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson· International Conference on...· 44 citations· ⚡5
It is demonstrated that linker-free PROTACs can outperform traditional designs, marking a paradigm shift in PROTAC development for targeted protein degradation.
Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or seque...
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.