The single-cell literature catalogs cell states as validated marker-gene programs — a sparse, compositional prior. Conventional embedding methods do not leverage this prior and learn cell-state structure de novo from the expression matrix, producing dense dimensions needing post-hoc interpretation and batch correction. Here we show coding agents can author single-cell embedding models directly from the literature. Given a scenario that focuses this literature lens on a chosen biological subdomain, the agent edits a structured Python template, curating named, literature-cited gene programs and composing them into axes, without a gene-set database, training, or sight of the data. Across mouse and human tissues these zero-shot embeddings are competitive in biological quality with conventional, foundation-model, and program-informed baselines, batch-robust by construction and reproducible across runs, complementing data-driven embeddings. Because each dimension is a named, cited gene program, the embedding is interpretable and auditable, and its composable axes can be steered into a developmental tree.
Niklas Brunn, S. M. Krißmer, Maximilian Frosch et al.· bioRxiv· 0 citations
Accurate identification of CRISPR-Cas9 off-target sites is essential for the safety assessment of genome-editing-based therapies. While numerous in silico prediction tools have been developed, their comparative performance and practical utility in preclinical workflows remain incompletely defined. We performed a systematic benchmarking of 14 in silico CRISPR-Cas9 off-target prediction tools, including both standard approaches and machine learning-based models. The analysis was based on a curated dataset derived from the CRISPRoffT database, comprising 3,827 deep-sequenced genomic sites across 26 guide RNA/Cas9 combinations in human cells. Sites with indel frequencies ≥0.1% were operationally defined as true off-targets. We evaluated tool performance using score distributions, correlation with indel frequencies, precision-recall characteristics, recall among top-ranked candidate sites, and the effect of combining tools. All tools assigned higher scores to true off-target sites compared with nontarget sites, although substantial overlap between classes was observed. Correlation between prediction scores and indel frequencies was weak to moderate, indicating limited ability to predict editing magnitude. Precision-recall performance was moderate across all tools, reflecting inherent trade-offs between sensitivity and specificity. Recall increased with the number of predicted sites considered, reaching approximately 77% among the top 500 and up to 83% among the top 1,250 sites, but leaving a substantial fraction of true off-targets undetected. Combining tools yielded only modest improvements. Current in silico tools enable prioritization of CRISPR-Cas9 off-target candidates but remain limited in their ability to comprehensively identify and quantitatively predict off-target activity. Our findings highlight the importance of considering both ranking performance and candidate site coverage and support the use of combined computational and experimental strategies for robust off-target assessment in preclinical gene editing workflows.
M. M. Kaufmann, Maren Hackenberg, William Jobson Pargeter et al.· Human Gene Therapy· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.