While foundation models have been shown to learn biological representations from large transcriptomic atlases, it remained unknown whether proteomics data allow the same. We here therefore introduce OmicsFM, a modality-agnostic transformer pretrained through masked abundance reconstruction on an unprecedented proteomics data corpus of 48,837 quality-filtered proteomics profiles from 1,397 reprocessed PRIDE projects. Interestingly, despite training on 14- to 93-fold fewer profiles than matched bulk- and single-cell transcriptomic models, respectively, our proteomics model rivals both. On held-out projects, OmicsFM attention networks recovered more molecular relationships than co-expression methods and existing single-cell foundation models across nine reference databases that reveal pathway-level organization. Sample-level embeddings preserved biological structure across independent studies, and its representations transferred successfully to cell-type classification, gene-essentiality prediction, and perturbation-response prediction, while consistently outperforming task-specific models. Moreover, our results show that proteomics and transcriptomics representations capture complementary biology. OmicsFM thus firmly establishes the possibility of training highly performant proteomics-based foundation models, and their importance in modelling and uncovering fundamental biology.
Sander Heyndrickx, R. Gabriels, Harikrishnan Ramadasan et al.· bioRxiv· 0 citations
Short The SwissProt database contains a stable 20,418 human protein-coding genes and 42,541 human protein sequences. Ribo-Seq suggests about 7,000 additional, non-canonical Open Reading Frames (ORFs) are present in humans, though only a few of them are confirmed by Mass Spectrometry (MS). Detecting these proteins requires extensive database searches, increasing computational load and inflating False Discovery Rates (FDR). Using the ionbot search engine with the OpenProt database allows for reliable detection of non-canonical proteins while controlling FDR. Ionbot surpasses the Trans-Proteomics Pipeline (TPP) in reproducibility, identifying more peptides and proteins supported by multiple spectra. In addition, open modification searches yield better PSMs compared to closed searches. This work highlights the importance of employing cutting-edge search engines in non-canonical protein research, as well as the value of open modification search in correcting errors in non-canonical protein detection. Long Background The SwissProt database reports a quite stable 20,418 human protein-coding genes and 42,541 human protein sequences, figures that have remained stable. New techniques like Ribo-Seq indicate that approximately 7,000 additional, non-canonical Open Reading Frames (ORFs) are translated in humans, few of which have been confirmed by Mass Spectrometry (MS). Detecting these non-canonical proteins requires comprehensive database searches, which increase computational load and False Discovery Rate (FDR). Here, we use the open search engine ionbot in combination with the OpenProt proteogenomics database to reproducibly detect non-canonical proteins while maintaining a well-controlled FDR. Results Compared to the current gold standard, the Trans-Proteomics Pipeline (TPP), ionbot shows higher reproducibility, with a higher number of peptides and proteins supported by multiple spectra, and across multiple samples. We observe that PSMs from the open modification search against OpenProt have higher fragment ion intensity correlation compared to PSMs obtained from the closed search, or by only searching canonical proteins. Conclusions In this work, we show the potential for open modification searching to correct potential mistakes in non-canonical proteins detection by preventing modified canonical peptides or variants from being incorrectly identified as non-canonical peptides. We also highlight the importance of assessing the FDR of non-canonical identifications separately from canonical ones, as global FDR calculations are biased by the scarcity of non-canonical identifications in each dataset.
V. Vasylieva, Enrico Massignani, Tine Claeys et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.