Abstract Many real-world problems are naturally modeled as heterogeneous graphs, where nodes and edges represent multiple types of entities and relations. Existing learning models for heterogeneous graph representation usually depend on the computation of specific, user-defined heterogeneous paths, or on the application of large, and often non-scalable, deep neural network architectures. We propose Het-node2vec, an extension of the node2vec algorithm, designed to embed heterogeneous graphs by capturing the topological and structural characteristics of the graph and the semantic information underlying the different types of nodes and edges; this is performed by introducing a simple stochastic node-type switching strategy in second-order random walk processes. Empirical results on synthetic graphs, as well as on benchmark and real-world biomedical graphs, show that Het-node2vec achieves comparable or superior performance to state-of-the-art methods for heterogeneous graphs in node label prediction tasks.
Mauricio Soto-Gomez, Carlos Cano, Justin Reese et al.· Scientific Reports· 0 citations
Large Language Models (LLMs) have transformed protein engineering by capturing complex sequence patterns from large datasets, enabling applications such as structure prediction and functional annotation. Finenzyme applies conditional transfer learning to generate biologically plausible enzyme sequences conditioned on Enzyme Commission (EC) numbers.
In this work, we extended Finenzyme with an in silico selection pipeline that first identifies generated sequences most likely to preserve or enhance the functional characteristics of specific EC categories and then evaluates them through molecular dynamics (MD) simulations to assess their structural stability and conformational dynamics.
MD simulations of 236 Finenzyme-generated enzymes across four EC classes (59
µ
s total simulation time) confirmed high structural stability. Across all enzyme classes, 74-95% of the models maintained stable tertiary structures and correct folding throughout the trajectories, with 195 out of 236 structures (82.6%) exhibiting sustained stability.
By combining conditional pre-trained language model fine-tuning with dynamic structural evaluation, our framework advances beyond static sequence-based predictions to address the structural and functional dimensions of enzyme behavior, key aspects for both biomedical and industrial applications.
E. M. Fassi, M. Nicolini, Emanuele Saitto et al.· Frontiers in Artificial Inte...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.