Skip to content
#protein folding Dataset Open access

RolyPoly-tk Data Bundle: reference profiles and sequence sets for virus analyses

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

RolyPoly data bundle Date / Bundle Version: 2026-09-03 This deposit contains the custom built and reference (3rd party) data required for all functionality of RolyPoly-tk, an RNA-Virus analysis tool kit. This includes external raw databases, bundled for convinenve, and custom, purpose built datasets only released and tracked here. Note that some of the public databases bundled have been processed and differ from the original (e.g. NCBI ribovirus sequences were masked for low complexity regions). Included here are the HMMer/mmseqs/cmscan profile collections, and all the various supporting sequence sets (rRNA collection, taxonomy mapping, adapter sequences, viral proteins with known functions) used by different commands - the marker protein and nucleic similarity based searches, raw NGS read filtering and processing, and genome annotation. Some of the buncdled datasets (e.g. the mitochodria, plastid, and tRNA sequence collections) are not currently exposed by rolypoly-tk codebase, and are should be treated as auxiliary collections prepared for experimental or future use. N.B. if you use any of the external or 3rd party datasets from this bundle, please cite them directly (bibliography details below). Code Compatibility: v0.7.19 Directory tree data/ ├── README.md ├── contam/ │ ├── adapters/ # adapter sequences used during read filtering │ ├── masking/ # viral nucleotide/protein masking references │ └── rrna/ # SILVA/NCBI rRNA masking references and mapping table ├── profiles/ │ ├── cm/ # pressed Rfam covariance-model database │ ├── hmmdbs/ # HMMER DBs, including split RVMT RdRp/RT profiles │ ├── mmseqs_dbs/ # matching MMseqs2 profile databases │ ├── genomad_rna_viral_markers_with_annotation.csv.gz │ ├── marker_features.tsv.gz │ ├── motif_metadata.json │ ├── NVPC_descriptions.csv.gz │ ├── pfam_rdrps_and_rts_profiles.tsv.gz │ ├── rvmt_motif_profiles.tsv.gz │ └── vfam.annotations.tsv.gz ├── reference_seqs/ │ ├── ncbi_virus/ # NCBI virus taxonomy DBs and RefSeq viral nucleotide DBs │ ├── RVMT/ # cleaned RVMT sequence DBs │ └── uniref/ # viral UniRef50 subset └── rrna/ # selected rRNA covariance models and cmscan audit tables Disclaimer Parts of this data archive were prepared from publicly available biological databases and reference resources. Before inclusion in RolyPoly, some of these resources were filtered, deduplicated, masked, compressed, converted into the formats required by and compatible with RolyPoly, or enriched with metadata and annotations. In some cases, only selected subsets or curated profiles were retained. It is uploaded here for versioning, sharing, and disclousre, and I can not vouch for using these data outside of rolypoly. We take and hold no responsibility for any outcome of any usage of this archive. Attribution note This archive includes third-party databases and derived products. We do not own these resources and we do not take responsibility for their completeness, accuracy, licensing, or any downstream use. Please respect the original licenses and data-use terms of the underlying sources. When using this archive, users should cite: RolyPoly itself The specific original third-party database(s) used for the relevant data subset Any software tools used to generate, index, or process the data, where applicable Note: rolypoly-tk commands and log files have a citation-reminder feature, that will try to provide you with the exact 3rd party tools and dbs used by the specific command you used. This can be turned on by changing the rpconfig.json file (see configuration). Third-party databases and citations Please cite the relevant original databases and software alongside RolyPoly when using these resources.A complete list of DOIs related to tools and databases is available in all_used_tools_dbs_citations.json Core reference databases RVMT Neri, U., Wolf, Y. I., Roux, S., Camargo, A. P., Lee, B., Kazlauskas, D., Chen, I.-M., Ivanova, N., Zeigler Allen, L., Paez-Espino, D., Bryant, D. A., Bhaya, D., Narrowe, A. B., Probst, A. J., Sczyrba, A., Kohler, A., Séguin, A., Shade, A., Campbell, B. J., Lindahl, B. D., Reese, B. K., Roque, B. M., DeRito, C., Averill, C., Cullen, D., Beck, D. A. C., Walsh, D. A., Ward, D. M., Wu, D., Eloe-Fadrosh, E., Brodie, E. L., Young, E. B., Lilleskov, E. A., Castillo, F. J., Martin, F. M., LeCleir, G. R., Attwood, G. T., Cadillo-Quiroz, H., Simon, H. M., Hewson, I., Grigoriev, I. V., Tiedje, J. M., Jansson, J. K., Lee, J., VanderGheynst, J. S., Dangl, J., Bowman, J. S., Blanchard, J. L., Bowen, J. L., Xu, J., Banfield, J. F., Deming, J. W., Kostka, J. E., Gladden, J. M., Rapp, J. Z., Sharpe, J., McMahon, K. D., Treseder, K. K., Bidle, K. D., Wrighton, K. C., Thamatrakoln, K., Nüsslein, K., Meredith, L. K., Ramirez, L., Buee, M., Huntemann, M., Kalyuzhnaya, M., Waldrop, M. P., Sullivan, M. B., Schrenk, M. O., Hess, M., Vega, M. A., O’Malley, M. A., Medina, M., Gilbert, N. E., Delherbe, N., Mason, O. U., Dijkstra, P., Chuckran, P. F., Baldrian, P., Constant, P., Stepanauskas, R., Daly, R. A., Lamendella, R., Gruninger, R. J., McKay, R. M., Hylander, S., Lebeis, S. L., Esser, S. P., Acinas, S. G., Wilhelm, S. S., Singer, S. W., Tringe, S. S., Woyke, T., Reddy, T. B. K., Bell, T. H., Mock, T., McAllister, T., Thiel, V., Denef, V. J., Liu, W.-T., Martens-Habbena, W., Allen Liu, X.-J., Cooper, Z. S., Wang, Z., Krupovic, M., Dolja, V. V., Kyrpides, N. C., and Koonin, E. V. (2022). The global virome of the ocean. Cell, 185(16), 2879–2894. https://doi.org/10.1016/j.cell.2022.08.023 Rfam Kalvari, I., Nawrocki, E. P., Ontiveros-Palacios, N., Argasinska, J., Lamkiewicz, K., Marz, M., Griffiths-Jones, S., Toffano-Nioche, C., Gautheret, D., Weinberg, Z., Rivas, E., Eddy, S. R., Finn, R. D., and Bateman, A. (2020). Rfam 14: expanded coverage of metagenomic, viral and microRNA families. Nucleic Acids Research, 49(D1), D192–D200. https://doi.org/10.1093/nar/gkaa1047 RefSeq O’Leary, N. A., Wright, M. W., Brister, J. R., Ciufo, S., Haddad, D., McVeigh, R., Rajput, B., Robbertse, B., Smith-White, B., Ako-Adjei, D., Astashyn, A., Badretdin, A., Bao, Y., Blinkova, O., Brover, V., Chetvernin, V., Choi, J., Cox, E., Ermolaeva, O., Farrell, C. M., Goldfarb, T., Gupta, T., Haft, D., Hatcher, E., Hlavina, W., Joardar, V. S., Kodali, V. K., Li, W., Maglott, D., Masterson, P., McGarvey, K. M., Murphy, M. R., O’Neill, K., Pujar, S., Rangwala, S. H., Rausch, D., Riddick, L. D., Schoch, C., Shkeda, A., Storz, S. S., Sun, H., Thibaud-Nissen, F., Tolstoy, I., Tully, R. E., Vatsan, A. R., Wallin, C., Webb, D., Wu, W., Landrum, M. J., Kimchi, A., Tatusova, T., and DiCuccio, M. (2015). Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Research, 44(D1), D733–D745. https://doi.org/10.1093%2Fnar%2Fgkv1189 SILVA Quast, C., Pruesse, E., Yilmaz, P., Gerken, J., Schweer, T., Yarza, P., Peplies, J., and Glöckner, F. O. (2012). The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Research, 41(D1), D590–D596. https://doi.org/10.1093/nar/gks1219 UniRef50 UniProt Consortium. UniProt Reference Clusters (UniRef). https://www.uniprot.org/help/uniref Profile and marker databases geNomad RNA viral markers Camargo, A. P., Roux, S., and others. geNomad: identification of mobile genetic elements and viruses from metagenomes. Nature Biotechnology, 2024. https://doi.org/10.1038/s41587-023-01953-y RdRp-Scan Charon, J., Buchmann, J. P., Sadiq, S., and Holmes, E. C. (2022). RdRp-scan: a bioinformatic resource to identify and annotate divergent RNA viruses in metagenomic sequence data. Virus Evolution, 8(2). https://doi.org/10.1093/ve/veac082 NeoRdRp v2.1 Sakaguchi, S., Urayama, S., Takaki, Y., Hirosuna, K., Wu, H., Suzuki, Y., Nunoura, T., Nakano, T., Nakagawa, S., and others. (2022). NeoRdRp: a comprehensive dataset for identifying RNA-dependent RNA polymerases of various RNA viruses from metatranscriptomic data. Microbes and Environments, 37(3). https://doi.org/10.1264/jsme2.ME22001 Pfam Mistry, J., Chuguransky, S., Williams, L., Qureshi, M., Salazar, G. A., Sonnhammer, E. L. L., Tosatto, S. C. E., Paladin, L., Raj, S., Richardson, L. J., Finn, R. D., and Bateman, A. (2020). Pfam: The protein families database in 2021. Nucleic Acids Research, 49(D1), D412–D419. https://doi.org/10.1093/nar/gkaa913 VFam / VOGDB The VOGDB / VFam resource. See the original VOGDB/VFam publications and associated database pages for citation details. https://doi.org/10.3390/v16081191 Software used to generate or index the data These tools were used in preparing the archive and should be cited where relevant: HMMER Eddy, S. R. (2011). Accelerated profile HMM searches. PLoS Computational Biology, 7(10), e1002195. https://doi.org/10.1371/journal.pcbi.1002195 MMseqs2 Steinegger, M. and Söding, J. (2017). MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature Biotechnology, 35(11), 1026–1028. https://doi.org/10.1038/nbt.3988 Infernal / cmsearch Nawrocki, E. P. and Eddy, S. R. (2013). Infernal 1.1: 100-fold faster RNA homology searches. Bioinformatics, 29(22), 2933–2942. https://doi.org/10.1093/bioinformatics/btt509 seqkit Shen, W., Sipos, B., and Zhao, L. (2024). SeqKit2: a Swiss army knife for sequence and alignment processing. iMeta, 3(3). https://doi.org/10.1002/imt2.191 BBMap Bushnell, B., Rood, J., and Singer, E. (2017). BBMerge – accurate paired shotgun read merging via overlap. PLOS ONE, 12(10), e0185056. https://doi.org/10.1371/journal.pone.0185056 tRNAscan-SE Lowe, T. M. and Eddy, S. R. (1997). tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence. Nucleic Acids Research

View source

Similar papers

#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6
#computer vision Book Open access Jul 2015

Understanding the affect of developers: theoretical background and guidelines for psychoempirical software engineering

This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.

D. Graziotin, Xiaofeng Wang, P. Abrahamsson · 56 citations · ⚡4
#machine learning Open access May 2017

What Influences the Speed of Prototyping? An Empirical Investigation of Twenty Software Startups

This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.

Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson · 44 citations · ⚡5
#protein folding Open access Sep 2026

Programmable design of functional proteins from natural language

Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or sequence constraints.

Fengyuan Dai, Shiyang You, Yudian Zhu et al. · 31 citations · ⚡3

Related blog posts

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

Google DeepMind Blog Nov 25, 2025

AlphaFold: Five years of impact

Explore how AlphaFold has accelerated science and fueled a global wave of biological discovery.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.