SPINET is introduced, which predicts sequences from molecular dynamics trajectories that uses cellular sheaves to represent residue interactions within each frame and recurrent units to integrate information across frames, then predicts all amino acids in a single pass.
Jens Lundsgaard, Colin Mikulski, Zhi-Xuan Yan et al.· 0 citations
A fully automated annotation framework for generating precise molecular descriptions at scale, such that the original molecule can be unambiguously reconstructed from the description alone and the proposed framework and dataset provide a reliable foundation for molecule--language alignment.
Recent advancements in image-based profiling techniques have enabled the collection of high-volume cell morphology data, allowing new molecular embedding models to learn from the experimental phenotypic perturbations of a molecule in a cell. Previously, we developed Molecule-Morphology Contrastive Pretraining (MoCoP),...
Jie Li, Kathryn E. Kirchoff, Dante A. Pertusi et al.· 0 citations
LapDDPM is introduced, a novel conditional Graph Diffusion Probabilistic Model designed for robust manifold learning and high-fidelity generation that significantly outperforms state-of-the-art baselines in distribution matching, manifold preservation, and downstream utility.
Lorenzo Bini, Stéphane Marchand-Maillet· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Protein modification requires navigating an immense sequence space, yet wet-lab validation remains low-throughput and costly. Although computational paradigms including protein language models (PLMs), large language models (LLMs), and LLM-based agents have shown promise in protein modification, their relative efficacy...
Ya-Wen Ouyang, Xin-Bo Zhang, Zi-Yuan Ma et al.· 0 citations
Protein language models are trained primarily with masked language modeling (MLM), which predicts masked amino-acid identities. Joint-embedding predictive architectures (JEPA) instead predict latent representations, but have not been applied to proteins.
ProteinJEPA supplements MLM with a cosine loss for predicting t...
Dan Ofer, Dafna Shahaf, Michal Linial· 0 citations
It is found that, at variance with networks not based on the attention mechanism, the lack of a first--order--like transition in the loss of the transformer produces a range of intermediate temperatures with good learning properties; this is true both for synthetic and natural protein sequences.
L. Ghiringhelli, A. Zambon, G. Tiana· arXiv.org· 0 citations
A hierarchical Bayesian multitask learning model that is applicable to the general multi-task binary classification learning problem where the model assumes a shared sparsity structure across different tasks is proposed and derived based on variational inference to approximate the posterior distribution.
Hao-Nan Zhu, Andre R. Goncalves, Car Reen Kok et al.· BioData Mining· 1 citation
Many molecular Transformers lack probabilistic latent variables for posterior inference and latent interpolation. We introduce STAR-VAE, a SELFIES-encoded, Transformer-based, AutoRegressive Variational AutoEncoder combining a bidirectional encoder with an autoregressive decoder pretrained on 79 million PubChem molecule...
Bum Chul Kwon, Ben Shapira, Moshiko Raboh et al.· 0 citations
Virtual screening aims to prioritize active compounds from large chemical libraries within a limited experimental budget. When applying Boltz-2 to virtual screening, a key challenge is how to use limited experimental data from the target assay to improve the prioritization of active compounds. We investigated whether f...
This work develops a systematic understanding of how physical patterns can alternatively be discovered directly from data by training a model without domain-specific priors, including any manually defined atomistic pairwise interactions, and finds that the model autonomously recovers key physical structure.
T. Kreiman, Yu-Tong Bai, Fadi Atieh et al.· 10 citations
STOMP is a powerful, robust multi-objective alignment algorithm that can meaningfully improve post-training in multiple domains and achieves or ties for the highest hypervolumes on 16/18 protein tasks and 5/6 natural language tasks.
Aadyot Bhatnagar, Peter Mørch Groth, Sebastian Ibarraran et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.