This work presents TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways, and introduces new instruction tasks for capabilities that current corpora barely cover.
Hyunjin Seo, Hyeon Hwang, Gyubok Lee et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.