Skip to content

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

Jul 2026 · arXiv.org · Vol abs/2607.08803 · 0 citations · 96 references
Computer Science Biology

TL;DR

This work presents TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways, and introduces new instruction tasks for capabilities that current corpora barely cover.

Abstract

The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-cell atlases, and pathway databases, are scattered across heterogeneous formats and remain unorganized into a cohesive corpus for language model training. We present TheBioCollection, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways. Beyond consolidating existing data, TheBioCollection enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover. We pair the corpus with TheBioCollection-Eval, a matched suite probing recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings. Holding the base Gravity-16B-A3B architecture fixed, training on TheBioCollection more than doubles its overall score on TheBioCollection-Eval with gains in every domain, while leaving general linguistic ability nearly intact.

View source

Similar papers

Open access Aug 2026

Protein language models and the long tail of functional diversity

It is found that many GigaRef singletons belong to a cluster under alternative parameter settings, suggesting that genomic and metagenomic datasets may require dataset-specific clustering configurations, and it is shown that singletons share mutual information with clustered sequences, making them learnable by PLMs and useful for training.

R. Vinod, Samir Char, Ava A. Amini et al. · 0 citations
Review Jul 2026

Genomic language models (gLMs): Emerging applications, challenges, and future directions in computational genomics.

Across the included studies, stronger evidence for gLM utility was generally associated with biologically informed or task-aligned model design, including evolutionary alignments, motif-aware objectives, long-context architectures, RNA structural priors, population-aware representations, and domain-specific pretraining.

Mahinaz A. Mashhour, M. A. Wahed, Mai S. Mabrouk · 0 citations
#protein folding Review Open access Aug 2026

Large language models in bioinformatics: a comprehensive survey

This survey reviews the basic principles of LLMs and summarizes representative applications in gene and genome sequence analysis, protein structure and function prediction, and drug design, including virtual screening and personalized medicine.

Zhi-Gang Meng, Zhi-Kai Yang, Mingming Zhu et al. · 0 citations
Open access Jul 2026

∑0–EvoCell: An AI-Native Ontology that Unifies Evolutionary and Cell Biology in Latent Space

Foundation models for biology achieve impressive pattern recognition on molecular sequences and single-cell transcriptomics, yet they fail to outperform simple linear baselines for predicting genetic perturbation effects, exposing a gap between statistical correlation and mechanistic understanding. This gap is compounded by an interface problem: biological knowledge lives in human-readable formats (FASTA, SBML, ontology triples) that must be lossily re-encoded before a neural network can reason about them. Here we introduce the Evolutionary Cell Ontology (ECO), an AI-native formal language built on the Σ0 substrate that represents biological knowledge directly as vector-encoded relational graphs. ECO uses 16 structural operators that serve simultaneously as knowledge glyphs, tensor operations, and—critically—carry a dual semantics spanning both evolutionary and cellular timescales, so the same operator that denotes speciation at the phylogenetic scale denotes irreversible APC/C commitment at the cell-cycle scale. We define a Latent Space Communication Protocol (LSCP) that maps ECO graphs into the residual stream of large language models, enabling systematic auditing of the biological knowledge a model actually contains. We illustrate ECO across three domains using controlled simulations: (i) globin protein-family evolution across roughly 1.5 billion years, where ECO epistatic attention recovers long-range coevolutionary couplings and ancestral-state reconstruction reaches 82–91% accuracy graded by conservation class; (ii) mammalian cell-cycle dynamics, where a CDK–cyclin attention graph with GATE checkpoints and FUSE commitment nodes reproduces two full oscillatory cycles with a four-attractor phase portrait emerging without explicit programming; and (iii) latent-space alignment, where ECO embedding distance tracks divergence across 60 protein families (Pearson r = 0.50) and an illustrative LSCP audit projects that relational, multi-scale concepts are encoded far more weakly than sequence-level facts. ECO replaces the human-readability constraint with an AI-processing constraint and, in doing so, turns the opacity of foundation models into a measurable, navigable coverage map.

Lurong Pan · 0 citations
Open access Sep 2026

LAMBDA: a prophage detection benchmark for genomic language models.

Transformer-based genomic sequence models represent an emerging frontier in computational biology. Yet, their embeddings have not yet shown the same level of predictive power as natural and protein language models, highlighting a gap between current implementations and theoretical promise. Existing benchmarks for DNA language models primarily focus on classifying regulatory elements in eukaryotic genomes, leaving open the fundamental question of whether these models learn sequence-level features across whole genomes. We introduce LAMBDA, a benchmark designed to rigorously evaluate genome language model embeddings through phage-bacteria sequence discrimination across four categories of increasing complexity: probing tasks, fine-tuning assessments, diagnostic tests, and genome-wide prophage detection. Our comprehensive analysis of current genomic language models provides insight into the importance of training data selection relative to model size, the need for domain-specific training, and the capabilities and limitations of genomic language models for detecting prophage sequences. This benchmark represents a challenging genomic annotation task in the bacterial domain and addresses a key computational problem with direct relevance to microbiology and medicine.

LeAnn M. Lindsey, Nicole L. Pershing, K. Dufault-Thompson et al. · 0 citations
#machine learning Preprint Aug 2026

Task- and dataset-specific information in protein language models

Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs). By consensus, embeddings from the models'last layers are used, while the models'internal behavior remains poorly understood. We analyzed 13 PLMs across 15 DTs and 9 datasets to assess the value of embeddings from intermediate PLM layers. We trained probe models on embeddings from each layer, compared their performance, and showed that the last layers of PLMs rarely produced embeddings that led to the best results on downstream tasks. Furthermore, we identified a connection between how models learn a certain DT and the similarity between that DT and the pre-training objective. For example, for residue-level downstream tasks, we observed a steady increase in performance across almost all PLM layers, which we attributed to their similarity to most PLMs'pre-training objectives. To allow the community to capitalize on our findings, we provide PLMSommelier, a Python package that automatically identifies the best PLM layer for a given DT with ~98% accuracy and creates a truncated model using only the early layers up to the best-performing layer. This will help users save time and memory during inference and yield better predictive performance.

R. Joeres, Ilya S. Senatorov, A. Kolchina et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.