Backcross: near-isogenic line characterisation from SNP genotypes
Abstract
Added Released together with progeny-selector (https://github.com/piercetaylor/progeny-selector/releases/tag/v0.1.0); both implement input data contract 1.12.0. Each GitHub Release carries backcross- -site.zip, the built site with a relative base that serves from any static folder (with a one-line server instruction, since Chrome and Edge do not run module scripts from file://), the CLI bundle, LICENSE and CITATION.cff; and the CLI bundle on its own. The release body is that version's CHANGELOG section plus the contract version it implements. A user guide for breeders (docs/user-guide.md) with a validation-status section, and a docs index (docs/README.md). CITATION.cff (validated in CI with cffconvert), a README citation section, and a sibling-tool section cross-linking progeny-selector. Every analysis CSV (summary, QC, segments, targets, pairwise, discordant markers) gains three trailing columns after crop: tool (backcross), tool_version and tool_commit (g + seven hex, -dirty when built from an uncommitted tree, NA when git was unavailable), and the HTML report a Software row and a generator meta tag, so a result names the release that produced it (docs/adr/0030). brapi-callsets.csv is unchanged. A single-file CLI bundle, dist-cli/backcross-cli.mjs, built by npm run build:cli with esbuild and attached to every GitHub Release; it needs only Node 22.19 and takes the same subcommands and options as node src/cli.ts (docs/adr/0030). A code of conduct (Contributor Covenant 2.1), a security policy, issue and pull-request templates and Dependabot configuration. qc.csv: one row per sample in manifest order, parents included, with missing, heterozygosity and nonparental rates and |-joined qc_flags, from the new CLI qc subcommand and an Export screen button (docs/adr/0029). The per-line summary CSV gains max_marker_coverage_bp and max_marker_coverage_cm, the resolved RPP maximum marker coverage the weighted estimators were computed with, before token_profile (docs/adr/0028). ADR 0006 gains an amendment correcting its Flapjack comparability claims: interior weighting follows Flapjack, totals differ at chromosome ends by design. Contract 1.12.0 (docs/adr/0027): the soybean scheme also reads the SoyBase / LIS Data Store names of Williams 82 (glyma.Wm82.gnmN.Gm01, glyma.Wm82.gnm5.Chr01) and the V1.1 spelling GLYMAchr_01 as Gm01..Gm20; every spelling read before maps as before, and RefSeq accessions, other cultivars and Data Store scaffolds are still kept as written. Contract 1.11.0 (docs/adr/0026): every text input (the genotype file plain, gzip or bgzip, samples.csv, markers.csv, and a custom token profile) must be UTF-8. A byte sequence that is not valid UTF-8 (a Latin-1 or Windows-1252 é saved as 0xE9, an overlong encoding, an encoded surrogate, a sequence cut short at the end of the file) is now the error text.invalid_utf8, naming the file, the line and the byte position (samples.csv line 4: not valid UTF-8 (byte 0xE9 at position 24)); before, backcross replaced it with U+FFFD without a word, so a sample_id could silently fail to match. Save such files as "CSV UTF-8". A byte-order mark is still ignored, and non-ASCII ids that are valid UTF-8 match across files. Contract 1.10.0 (docs/adr/0025): the VCF GT grammar is stated. A sign, an exponent, an underscore, a space, a leading zero, an empty side (/0, 0/), three or more alleles, the VCF 4.4 leading phase indicator (|0|1), or an allele index above 127 or outside REF,ALT is now the error genotypes.invalid_gt, naming the line and the value; before, -5/0 was stored as allele 251, 1e1/0 read as 10 and /0 as REF. Haploid calls, ./., ., ./1, an empty GT (missing) and multiallelic indices are read as before. CLI subcommands compare and discordant run one pairwise comparison (algorithm 5) between --a and --b and write the pairwise summary CSV and the discordant-marker CSV respectively; --mode selects informative (default) or all. A CI job r-reader generates every CLI CSV table (six since qc) from the synthetic fixture and reads them back with readr under explicit column types (scripts/read_exports.R), modelled on progeny-selector's r-reader job. A crop scheme for sunflower (contract 1.9.0, docs/adr/0024). The Crop select and --crop now offer sunflower, whose canonical names are bare 1..17: Ha412HOChr01 (HA412-HOv2.0) and HanXRQChr01 (HanXRQr2.0-SUNRISE) are both read as 1, because the two assemblies share the LG1-LG17 numbering (XRQ and HA412-HO v1 by their map anchoring, HA412-HOv2 by a whole-genome alignment against XRQ; docs/adr/0024 gives the evidence and its weakest link), and chr1, chromosome-1, 01 and 1 are read too. Unplaced scaffolds (HanXRQChr00c001), organelles, accessions (NC_035433.2, CM007890.2), the LG prefix and HA412-HO v1.1 names (Ha1) are kept as written. A case under contract/cases/ pins the spellings, including two it leaves as written. The shared input contract is version 1.8.0: a genotype cell pairing two of N, -, ., in any order and in all three spellings (N/N, N-, ./N, .., -|N), is read as missing in HapMap and in nucleotide-mode wide CSV. Both tools already did this and neither changes; the contract now says so, with four cases under contract/cases/ and a mirror in progeny-selector (docs/adr/0023). A pair with X that is not itself a missing token (X/X, XN) stays an error, and under a base: none token profile (dart, axiom, kasp) such a cell is genotypes.unknown_cell unless the profile lists it, as axiom lists --. A wide CSV holding such a cell that is not itself a missing token is read as nucleotide by auto detection. The demo link takes a crop: ?demo=synthetic&crop=pea loads the demo under that crop scheme and leaves the Crop select on it, and "Copy link to this demo" adds &crop= for any crop but soybean, so the plain ?demo=synthetic link is unchanged. An unknown crop in the address shows an alert naming the built-in crops and loads nothing. The Crop select also starts at the loaded dataset's crop when you return to Upload (docs/adr/0017, amendment of 2026-09-26). Crop schemes for cowpea, pea and peanut (contract 1.7.0, docs/adr/0022). The Crop select and --crop now offer cowpea (Vu01..Vu11, which also reads NCBI's Vu01(old4) names in their eleven exact pairings), pea (chr1LG6..chr7LG7, which reads chr4 and 4LG4 but keeps a bare 4 as written, because pea tables use bare digits for two numberings that disagree) and peanut (Arahy.01..Arahy.20, where A01..A10, B01..B10 and the Aradu./Araip. names are kept as written). Sunflower stayed deferred then and followed as contract 1.9.0 above. Three cases under contract/cases/ pin each scheme's spellings, including one it leaves as written. The shared input contract is version 1.6.0: a genotype cell pairing one of A, C, G, T with one of N, -, ., in either order and in all three spellings (AN, -A, A/N, .|G), is read as missing in HapMap and in nucleotide-mode wide CSV. Both tools already did this and neither changes; the contract now says so, with four cases under contract/cases/ and a mirror in progeny-selector (docs/adr/0021). AX and A? stay errors, and under a base: none token profile (dart, axiom, kasp) the pair is still genotypes.unknown_cell. A pair of two missing characters (N/N, ..) was left undefined by 1.6.0 and is defined by 1.8.0 above. Crop selector (contract 1.5.0, docs/adr/0020): soybean (default), maize, rice, sorghum, wheat, barley, oat, common bean and cotton chromosome schemes; a chosen crop normalises and orders chromosome names by that crop's convention, so a maize chr1 is no longer displayed as Gm01. Every CSV export gains a trailing crop column and the report a Crop row. Token profiles (contract 1.4.0, docs/adr/0019): the Upload screen, the CLI (--profile) and the sibling read HapMap and wide-CSV cells under a named vocabulary, tassel, soybase-report (H heterozygous, U missing), dart (0/1/2/-), axiom (AA/AB/BB, NoCall, or 0/1/2) or kasp (X:X/X:Y/?), or a JSON file of the same shape. Every CSV export gains a trailing token_profile column and the report a Token profile row; readers by position keep their columns. A profile with a VCF, a BrAPI source or a wide CSV that is coded A/B/H is an error, as is an H at a marker with an indel allele; a profile file is always recorded as custom: . The CLI names the file in a profile read error. Nine new cases under contract/cases/. A demo dataset. The Upload screen's "Try the demo dataset" button loads the synthetic test fixture (six generated lines, 500 markers) that the site serves under demo/synthetic/, through the same load path as picked files. "Copy link to this demo" copies a link that opens the page at ?demo=synthetic, which loads the demo on arrival and opens the summary; an unknown demo value is reported in the alert. The demo data are synthetic, never real genotypes (docs/adr/0017). Accessibility review (M3): focus order verified in Chromium and Firefox from the skip link, through the rail, into every screen; a Lines row reached by keyboard is never hidden under the sticky header (WCAG 2.4.11); every status carries a text or ARIA cue beside its colour or weight; axe-core reports no WCAG 2.0/2.1 A or AA violation on any of the six screens, including the BrAPI form, a load in flight with its Cancel button, its error, the empty-filter table, the chromosome overview, a region error and the sort list open; and Lighthouse accessibility is above 90 on the built landing page, run by npm run a11y:lighthouse locally and in CI. Visible changes: the "input coding reference" link says it opens in a new tab; the genotype canvases are exposed as images with their descriptions; the rail's steps are a toolbar inside the navigation landmark; cancelling a BrAPI load moves focus to the Upload heading instead of dropping it; and printing hides the rail, the action bars, the genotype toolbar and the hover panel, unrolls the Lines table, and keeps the legend swatches' colours and textures. Load a variant set from a BrAPI v2.1 s