This work presents TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways, and introduces new instruction tasks for capabilities that current corpora barely cover.
Hyunjin Seo, Hyeon Hwang, Gyubok Lee et al.· arXiv.org· 0 citations
This work empirically finds that transferability depends on architecture-models with fully shared transformer backbone and a unified visual encoder exhibit consistent cross-task transfer, while loosely coupled designs show little or none.
Jiwon Kang, Heeji Yoon, Jaewoo Jung et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.