Skip to content

Category

bioinformatics

54 papers

#machine learning Preprint Sep 2026

SPINET: Sheaf Protein Inverse Folding Network

SPINET is introduced, which predicts sequences from molecular dynamics trajectories that uses cellular sheaves to represent residue interactions within each frame and recurrent units to integrate information across frames, then predicts all amino acids in a single pass.

Jens Lundsgaard, Colin Mikulski, Zhi-Xuan Yan et al. · 0 citations
#artificial intelligence Preprint Feb 2026

MolLangData: A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method

A fully automated annotation framework for generating precise molecular descriptions at scale, such that the original molecule can be unambiguously reconstructed from the description alone and the proposed framework and dataset provide a reliable foundation for molecule--language alignment.

Feiyang Cai, G. He, Yi Hu et al. · 0 citations
#machine learning Preprint Open access Sep 2026

Improving Molecular-Morphology Contrastive Pretraining using Deep-Learning-based Morphology Profiles

Recent advancements in image-based profiling techniques have enabled the collection of high-volume cell morphology data, allowing new molecular embedding models to learn from the experimental phenotypic perturbations of a molecule in a cell. Previously, we developed Molecule-Morphology Contrastive Pretraining (MoCoP),...

Jie Li, Kathryn E. Kirchoff, Dante A. Pertusi et al. · 0 citations
#artificial intelligence Preprint Jun 2025

LapDDPM: Spectral Perturbation Diffusion for Robust Single-Cell Manifold Generation

LapDDPM is introduced, a novel conditional Graph Diffusion Probabilistic Model designed for robust manifold learning and high-fidelity generation that significantly outperforms state-of-the-art baselines in distribution matching, manifold preservation, and downstream utility.

Lorenzo Bini, Stéphane Marchand-Maillet · 0 citations
#artificial intelligence Preprint Sep 2026

PFArena: Benchmarking Language Models for Protein Modification

Protein modification requires navigating an immense sequence space, yet wet-lab validation remains low-throughput and costly. Although computational paradigms including protein language models (PLMs), large language models (LLMs), and LLM-based agents have shown promise in protein modification, their relative efficacy...

Ya-Wen Ouyang, Xin-Bo Zhang, Zi-Yuan Ma et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

ProteinJEPA: Latent prediction improves protein language model pretraining

Protein language models are trained primarily with masked language modeling (MLM), which predicts masked amino-acid identities. Joint-embedding predictive architectures (JEPA) instead predict latent representations, but have not been applied to proteins. ProteinJEPA supplements MLM with a cosine loss for predicting t...

Dan Ofer, Dafna Shahaf, Michal Linial · 0 citations

Sampling at intermediate temperatures is optimal for training large language models in protein structure prediction

It is found that, at variance with networks not based on the attention mechanism, the lack of a first--order--like transition in the loss of the transformer produces a range of intermediate temperatures with good learning properties; this is true both for synthetic and natural protein sequences.

L. Ghiringhelli, A. Zambon, G. Tiana · 0 citations
#machine learning Open access Feb 2025

Hierarchical sparse Bayesian multitask learning for disease prediction in pooled microbiome studies

A hierarchical Bayesian multitask learning model that is applicable to the general multi-task binary classification learning problem where the model assumes a shared sparsity structure across different tasks is proposed and derived based on variational inference to approximate the posterior distribution.

Hao-Nan Zhu, Andre R. Goncalves, Car Reen Kok et al. · 1 citation
#artificial intelligence Preprint Open access Sep 2026

STAR-VAE: A Scalable Latent-Variable Transformer for Controllable Molecular Generation

Many molecular Transformers lack probabilistic latent variables for posterior inference and latent interpolation. We introduce STAR-VAE, a SELFIES-encoded, Transformer-based, AutoRegressive Variational AutoEncoder combining a bidirectional encoder with an autoregressive decoder pretrained on 79 million PubChem molecule...

Bum Chul Kwon, Ben Shapira, Moshiko Raboh et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Adapting Boltz-2 with limited experimental activity data improves early enrichment in virtual screening

Virtual screening aims to prioritize active compounds from large chemical libraries within a limited experimental budget. When applying Boltz-2 to virtual screening, a key challenge is how to use limited experimental data from the target assay to improve the prioritization of active compounds. We investigated whether f...

Kairi Furui, Masahito Ohue · 0 citations
#machine learning Preprint Oct 2025

Transformers Discover Molecular Structure Without Graph Priors

This work develops a systematic understanding of how physical patterns can alternatively be discovered directly from data by training a model without domain-specific priors, including any manually defined atomistic pairwise interactions, and finds that the model autonomously recovers key physical structure.

T. Kreiman, Yu-Tong Bai, Fadi Atieh et al. · 10 citations

Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebysheff Scalarization

STOMP is a powerful, robust multi-objective alignment algorithm that can meaningfully improve post-training in multiple domains and achieves or ties for the highest hypervolumes on 16/18 protein tasks and 5/6 natural language tasks.

Aadyot Bhatnagar, Peter Mørch Groth, Sebastian Ibarraran et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.