Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Unify learns cellular evolution with universal multimodal embeddings

Integrating single-cell RNA-sequencing (scRNA-seq) data across species is hindered by evolutionary divergence, technical batch effects, and the reliance on one-to-one orthologs. Here, we present Unify, a transfer learning methodology that learns universal cell embeddings by defining functionally coherent, multi-modal macrogenes. This is achieved by combining RNA expression with embeddings from protein language models and general-purpose language models. Unify transcends species boundaries, enabling cross-species comparisons beyond strict gene-level homology. Unify corrects batch effects while preserving conserved biological signals across vast evolutionary distances and enables more accurate prediction of perturbation responses across species, such as from mouse to human. Applied to species separated by over 700 million years, Unify reconstructs more accurate multi-species cell-type evolutionary trees and uncovers convergent gene programs. Together, these results establish Unify as a powerful method for comparative single-cell genomics and evolutionary biology. Integrating single-cell RNA-sequencing (scRNA-seq) data across species is still technically challenging. Here, the authors report a transfer learning framework designed to integrate scRNA-seq data across species by combining RNA expression with embeddings from protein language models and general-purpose language models.

Hua-Wen Zhong, Wenkai Han, Guoxin Cui et al. · 0 citations
Open access Jul 2026

ProtSyntax: a protein large language model for decoding post-translational modification syntax and function

Post-translational modifications (PTMs) expand protein function by encoding context-dependent regulatory states, and their dysregulation contributes to cancer, neurodegeneration and metabolic disease. However, existing methods treat PTMs as independent residue labels, limiting their ability to distinguish contextually permissible sites, model crosstalk and infer functional consequences. Here we introduce ProtSyntax, a PTM-aware foundation protein language model combining protein-aware positional encoding, bidirectional state-space propagation, geometry-constrained attention and adaptive multi-objective learning. This design integrates residue chemistry, motif order, long-range context and three-dimensional microenvironments while coupling PTM recognition to enzyme function. Across 40 PTM-site benchmarks, ProtSyntax exceeded the strongest baselines in mean MCC and AP by 12.66% and 10.67%. ProtSyntax also recovered masked PTM types and sites, rejected structural decoys, generalized to data-scarce modifications, reconstructed crosstalk and linked PTM perturbations to enzyme kinetics. Applications to pathogenic variants, biomolecular condensates and disease-associated PTM landscapes demonstrate its potential to decode the regulatory language of the modified proteome.

Yiyu Lin, Jiahui Wu, You Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.