Recent instrumental and computational innovations in mass-spectrometry-based proteomics offer new promise in biomarker discovery, thanks to unprecedented proteome coverage and depth. Data-independent acquisition (DIA) methods are very promising in this context as they allow improved proteome coverage, reduced missing value rates, and enhanced quantification precision. However, DIA methods also suffer from their own challenges, such as increased data complexity, cycle times, and background noise. In this work, we propose a sample-aware diaPASEF method optimization strategy for a timsTOF platform. Thorough method optimizations have first been conducted on standard HeLa lysates. Then, a ground-truth calibrated sample series, consisting of a range of UPS amounts spiked into a complex Arabidopsis background, was used to mimic differential analyses under controlled conditions. These benchmark experiments demonstrate clear benefits of using narrowPASEF for differential protein discovery. Finally, our strategy was applied to real use case biological samples to conduct a differential analysis of purified mouse astrocyte cells across two different conditions. narrowPASEF improved the proteome depth by 13%, considering proteins quantified with a coefficient of variation (CV) of <20%, and led to a 68% (435 vs 729) increase in differentially expressed proteins. These results provide an opportunity for a more precise and comprehensive analysis of the biological functions of biomarkers, offering a more profound understanding of the disease mechanisms. The benefits of our sample-aware narrowPASEF strategy demonstrated the most substantial impact on low-abundance proteins. Overall, these results show promise for more valuable and robust biomarker discoveries in the future.
While foundation models have been shown to learn biological representations from large transcriptomic atlases, it remained unknown whether proteomics data allow the same. We here therefore introduce OmicsFM, a modality-agnostic transformer pretrained through masked abundance reconstruction on an unprecedented proteomics data corpus of 48,837 quality-filtered proteomics profiles from 1,397 reprocessed PRIDE projects. Interestingly, despite training on 14- to 93-fold fewer profiles than matched bulk- and single-cell transcriptomic models, respectively, our proteomics model rivals both. On held-out projects, OmicsFM attention networks recovered more molecular relationships than co-expression methods and existing single-cell foundation models across nine reference databases that reveal pathway-level organization. Sample-level embeddings preserved biological structure across independent studies, and its representations transferred successfully to cell-type classification, gene-essentiality prediction, and perturbation-response prediction, while consistently outperforming task-specific models. Moreover, our results show that proteomics and transcriptomics representations capture complementary biology. OmicsFM thus firmly establishes the possibility of training highly performant proteomics-based foundation models, and their importance in modelling and uncovering fundamental biology.
Sander Heyndrickx, R. Gabriels, Harikrishnan Ramadasan et al.· bioRxiv· 0 citations
Short The SwissProt database contains a stable 20,418 human protein-coding genes and 42,541 human protein sequences. Ribo-Seq suggests about 7,000 additional, non-canonical Open Reading Frames (ORFs) are present in humans, though only a few of them are confirmed by Mass Spectrometry (MS). Detecting these proteins requires extensive database searches, increasing computational load and inflating False Discovery Rates (FDR). Using the ionbot search engine with the OpenProt database allows for reliable detection of non-canonical proteins while controlling FDR. Ionbot surpasses the Trans-Proteomics Pipeline (TPP) in reproducibility, identifying more peptides and proteins supported by multiple spectra. In addition, open modification searches yield better PSMs compared to closed searches. This work highlights the importance of employing cutting-edge search engines in non-canonical protein research, as well as the value of open modification search in correcting errors in non-canonical protein detection. Long Background The SwissProt database reports a quite stable 20,418 human protein-coding genes and 42,541 human protein sequences, figures that have remained stable. New techniques like Ribo-Seq indicate that approximately 7,000 additional, non-canonical Open Reading Frames (ORFs) are translated in humans, few of which have been confirmed by Mass Spectrometry (MS). Detecting these non-canonical proteins requires comprehensive database searches, which increase computational load and False Discovery Rate (FDR). Here, we use the open search engine ionbot in combination with the OpenProt proteogenomics database to reproducibly detect non-canonical proteins while maintaining a well-controlled FDR. Results Compared to the current gold standard, the Trans-Proteomics Pipeline (TPP), ionbot shows higher reproducibility, with a higher number of peptides and proteins supported by multiple spectra, and across multiple samples. We observe that PSMs from the open modification search against OpenProt have higher fragment ion intensity correlation compared to PSMs obtained from the closed search, or by only searching canonical proteins. Conclusions In this work, we show the potential for open modification searching to correct potential mistakes in non-canonical proteins detection by preventing modified canonical peptides or variants from being incorrectly identified as non-canonical peptides. We also highlight the importance of assessing the FDR of non-canonical identifications separately from canonical ones, as global FDR calculations are biased by the scarcity of non-canonical identifications in each dataset.
V. Vasylieva, Enrico Massignani, Tine Claeys et al.· bioRxiv· 0 citations
Post-translational modifications (PTMs) and genetic variants regulate protein function, signalling, and disease, but their interpretation requires integration of sequence annotations with structural, interaction, and biophysical context. Although resources such as Scop3P, UniProt, the Protein Data Bank, and AlphaFold provide extensive annotations and structural information, integrating these data into reproducible structure-aware analyses still requires custom scripting and manual coordination between multiple independent tools. To address this challenge, we developed Scop3P-Toolkit, an open-source executable analytical environment for interactive analysis of PTMs, mutations, and proteomics-derived peptides in their structural context. The toolkit integrates protein annotation retrieval with structural mapping, residue interaction network analysis, comparative structural analysis, and residue-level biophysical profiling within a unified framework. Experimentally supported phosphosites, phosphopeptides, and phosphoproteomics evidence are provided for human proteins through Scop3P, with optional integration of curated UniProt PTM annotations. UniProt-derived PTMs, sequence features, and genetic variants are available for proteins from any species, extending the framework beyond the human phosphoproteome. Scop3P-Toolkit supports structure-centric analyses including interpretation of PTMs and disease-associated variants, analysis of residue interaction networks and their rewiring across alternative conformations, structural localisation of peptides, and exploration of protein–protein, protein–ligand, and host–pathogen interfaces. Interactive visualisation links sequence annotations, three-dimensional structures, residue interaction networks, and biophysical profiles, enabling coordinated exploration across multiple molecular representations. The toolkit is distributed as Jupyter notebooks, browser-based Voilà applications, and a Galaxy interactive tool, providing transparent, accessible, and reproducible workflows for both computational and experimental researchers. By integrating biological annotation resources into executable, structure-aware workflows, Scop3P-Toolkit enables reproducible interpretation of PTMs, mutations, and proteomics data.
Adrián Díaz, Natalia Tichshenko, Boris Depoortere et al.· bioRxiv· 0 citations
A major update of Scop3P, a proteomics-informed knowledgebase that contextualizes human phosphorylation sites within integrated sequence, structural, biophysical, evolutionary, and mutational frameworks, and provides a scalable and provenance-aware resource for phosphosite interpretation, hypothesis generation, and data-driven modelling of phosphorylation-dependent regulation.
P. Ramasamy, Natalia Tichshenko, Adrián Díaz et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.