Skip to content
#edge computing Open access

From plant pangenome assemblies to operational graph references

Oct 2026 · Frontiers in Plant Science · 26 references
Genomics and Phylogenetic Studies

Abstract

Plant pangenomes have established that a single linear assembly cannot represent the sequence and structural diversity of a species. High-quality assemblies reveal presence-absence variation, copy-number change, inversions, translocations, transposable-element polymorphism, and deeply divergent haplotypes across crops and their wild relatives (Liu et al., 2020;Jayakodi et al., 2020;Walkowiak et al., 2020;Hufford et al., 2021;Zhou et al., 2022;Gaccione et al., 2026). Recent reviews have consequently emphasized graph construction, graph-aware genotyping, functional panomics, and the distinctive challenges posed by plant genomes (Schreiber et al., 2024;Bao and Weigel, 2025;Cochetel and Cantu, 2026). In parallel, standards initiatives have proposed common identifiers, metadata, quality control, evaluation, and data-sharing practices for plant pangenomes (Heuermann et al., 2025;Kopalli et al., 2025). A narrower infrastructure gap nevertheless remains: a collection of assemblies is not an operational pangenome, and a downloadable graph file is not necessarily a usable graph reference. I argue that mature plant pangenome projects should release an operational, versioned graph-reference package as a standard deliverable. Such a package combines graph topology and biological paths with projected annotations, coordinate mappings, task-specific indexes, validation datasets, reproducible workflows, release histories, and governance. This distinction shifts the discussion from whether graphs are useful, which is already well established, to what must accompany a graph before a research community can use it routinely.Assemblies describe diversity but leave each user to reconstruct relationships among genomes.Whole-genome alignment, repeat handling, variant normalization, path naming, and coordinate translation are consequential choices. Repeating them independently makes analyses inconsistent and unnecessarily expensive. Variant call files reduce part of this burden, but they remain anchored to a privileged reference and can represent nested, multiallelic, or overlapping structural variants awkwardly. Alleles and genes absent from that reference may be difficult to map, genotype, or annotate consistently.A sequence graph addresses the representation problem by encoding sequences as nodes, alternative adjacencies as edges, and chromosomes or haplotypes as paths (Paten et al., 2017).Representation alone, however, does not create an operational reference. Graphs intended for routine reuse should provide stable accession and path identifiers; the exact input assemblies and quality metrics; mappings to legacy coordinates; gene, repeat, and regulatory annotations on relevant paths; and documented rules for representing uncertainty, collapsed sequence, and homeologous regions. They should also include prebuilt indexes for maintained analysis workflows, a parameter-complete construction pipeline, and persistent versioned releases. An operational reference is therefore not a particular file format or software product but a reproducible package linking biological content to supported analytical tasks.RNA-seq data from diverse varieties or accessions are still commonly analyzed by aligning all reads to a single reference genome. Transcripts carrying divergent alleles, accession-specific exons, or genes absent from that reference may map less efficiently, be misassigned, or fail to map, biasing transcript detection, expression estimates, and allele-specific expression. Genotype-specific references improve transcript quantification in several crops, and spliced pangenome graphs can support haplotype-aware pantranscriptome analysis (Sibbesen et al., 2023;Cochetel and Cantu, 2026).The remaining obstacle is not simply the absence of algorithms. Graph-aware aligners cannot be used routinely when species-level graphs, projected transcript annotations, splice-aware indexes, path conventions, and coordinate mappings are unavailable or incompatible. A genomic graph distributed without these components may be technically valid yet unusable for transcriptomics.RNA-seq therefore provides a concrete test of operational readiness: a released graph should support an ordinary user in moving from reads to interpretable expression estimates through a documented and benchmarked workflow. The same principle applies to resequencing, chromatin profiling, epigenomics, and structural-variant genotyping.A mature pangenome should be supported by multiple quality-controlled, chromosome-scale assemblies, stable accession metadata, and adequate representation of the diversity for which the resource is intended. Its graph release should include interoperable topology and embedded paths, assembly and sample provenance, projected annotations, legacy-coordinate mappings, precomputed indexes, containerized workflows, tutorials, test data, and explicit computing requirements. Canonical biological content should remain separable from software-specific indexes so that communities can support evolving toolchains such as vg, minigraph, minigraph-cactus, PGGB, and ODGI (Garrison et al., 2018;Li et al., 2020;Guarracino et al., 2022;Hickey et al., 2024;Garrison et al., 2024).Validation should accompany publication rather than be deferred to downstream users. Releases should report graph growth as genomes are added, sequence and path completeness, topological complexity, mapping accuracy, reference bias, variant-genotyping performance, and computational cost. Held-out accessions or haplotypes are preferable because evaluation only on genomes used for construction overstates generalizability. Benchmarks should include repetitive and structurally complex regions and, where relevant, polyploid-specific tests of homeolog assignment, subgenome labels, phasing, and collapsed sequence (Kopalli et al., 2025).Graph composition also requires scrutiny. A panel dominated by elite or closely related material can reproduce sampling bias even when its graph is technically complete. Conversely, adding increasingly divergent genomes may improve representation while increasing topology, index size, and ambiguous mapping. Graph references should therefore state their sampling purpose and limitations. They can reduce biological reference bias while increasing computational and training demands; those trade-offs should be measured rather than obscured.PlantPan, maintained by the National Genomics Data Center and China National Center for Bioinformation, shows both the value and the limitations of centralized graph resources. It integrates chromosome-scale assemblies from 11 plant species with gene groups, genomic variation, synteny, functional annotations, resistance genes, transcription factors, pathways, and homology relationships. Its workflow aligns assemblies to a selected reference with NUCmer (Marçais et al., 2018), detects differences with SyRI (Goel et al., 2019), and incorporates 13 reported variant classes into graphs with VG (PlantPan, 2026). Downloadable graph resources were listed for nine of the 11 species at the time of consultation. This centralized catalog reduces duplicated whole-genome comparisons and connects nonreference sequences and structural alleles with biological information. Yet the existence of a graph file does not establish a complete graph-reference system. Pairwise discovery against a selected linear reference may retain reference-dependent representations, and graph topology alone does not guarantee stable path identifiers, projected annotations, coordinate mappings, precomputed task-specific indexes, benchmark datasets, release histories, or maintained mapping and genotyping workflows. PlantPan is therefore a valuable foundation and an instructive test case: centralized distribution is feasible, but operational readiness requires additional layers of documentation, indexing, validation, and maintenance.No single graph will serve every biological question. A graph optimized for short-read mapping may differ from one designed for comparative annotation, phylogenomics, or visualization. A practical model is a federated family of compatible resources: a stable species backbone, population-or breeding-pool graphs, and focused graphs for complex loci. Shared sample identifiers, path naming, metadata schemas, and coordinate mappings would permit interoperability without forcing all diversity into one monolithic object.Federation requires governance as well as formats. Each resource needs designated maintainers, persistent release identifiers, transparent procedures for correcting or withdrawing assemblies, compatibility rules, and retained prior versions. Funding agencies should treat these activities as research infrastructure; journals should require deposition of graph packages and build workflows when a pangenome is presented as reusable; and repositories should preserve canonical graphs, metadata, and supported indexes. Graph construction, benchmarking, annotation, and maintenance should be recognized as scholarly contributions.Access must also be practical. If graph use requires specialized hardware, undocumented pipelines, or transfer of terabyte-scale intermediates, it may widen participation gaps. Projects should offer downloadable canonical files, compact indexes, cloud-ready workflows, representative tutorials, and clear resource estimates. Training should connect graphs to familiar tasks such as read mapping, expression analysis, genome-wide association studies, and candidate-gene inspection.Plant genomics no longer needs another general demonstration that linear references are incomplete. It needs reusable infrastructure that allows alternative sequence and haplotype diversity to enter everyday analysis. Assemblies document diversity; a graph encodes relationships among alternative sequences; an operational graph-reference package adds the annotations, indexes, coordinate mappings, benchmarks, workflows, release history, and governance required for routine reuse.The proposed standard can begin modestly. Mature projects should release at least one documented graph with embedded assembly paths, projected annotations, legacy-coordinate mappings, supported indexes, a reproducible build, and held-out benchmarks. Communities can then maintain federated releases for distinct populations and analytical tasks. This proposal complements recent reviews and standards efforts by defining graph publication as an operational commitment rather than a one-time file deposit. Without that commitment, plant pangenomes may remain impressive catalogs used mainly by specialists. With it, their structural and haplotypic diversity can become a routine substrate for plant genomics and breeding.The author confirms being the sole contributor of this work and has approved it for publication.

Read PDF

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.