isoform-dominance: isoform-usage quantification and discrimination from bulk RNA-seq
Abstract
isoform-dominance 2.4.1 A maintenance release from a post-release audit of 2.4.0. Several fixes change an answer, because the code now does what the documentation already said or because 2.4.0's answer was wrong; they are listed first. The rest stop a traceback or a misleading message, or fix a record, a script or the documentation. The full list, with the measurement behind each fix, is the 2.4.1 section of the changelog. Behaviour that changes on upgrade identifiability --min-log2fc: an estimand past the linearisation limit counts as not resolved (exit 3), as paper.md and the 2.3.0 notes already said. 2.4.0 compared the figure with the threshold and ignored beyond_linear, so such a contrast could exit 0. effect_resolvable becomes false there; the figures and the beyond_linear flags in --json are unchanged. identifiability exits 1 on a --window longer than --read-length, and on a config that lists one transcript in two groups. identifiability exits 1 on a config gene_id Ensembl does not know. 2.4.0 judged the classes against an empty gene background and exited 0. An identifiability --inputs rerun under another grouping no longer puts the saved transcripts the config does not name into a gene background the saved run did not have. annotate exits 1 on a human symbol none of whose genes is on a reference chromosome, listing each gene of the name with its region. HLA-DRB3 at release 116 is such a symbol, and Ensembl 116's cdna.all holds 23 protein-coding ones. --gene-id takes one of the genes anyway, with a NOTE. The reference-chromosome and chrX/chrY rules now apply to human only: for another species every gene of the name is a candidate, and two or more stop with AmbiguousGene (--gene-id picks one). 2.4.0 could drop a zebrafish gene on chr23 for a same-name gene on a chromosome with a human name. extract counts one transcriptome indexed with and without Salmon's decoys as two indexes, so such a cohort is refused like any other mix (--allow-mixed-index combines it with a warning). A meta_info.json it cannot read is a warning, not a refusal or a traceback. extract writes its rows in quant.sf path order again, as 2.3.0 did. 2.4.0 sorted them by donor name; the two orders differ when one donor name is a prefix of another and the next character sorts before / (D1, D1-2, D1.5), and stats bootstraps the fold interval by row. Output: the identifiability header names the window, not k, and the per-class line separates the class's unique k-mers from its best transcript's stretch. The extract sidecar and saved inputs gain fields; files written by 2.4.0 still read, and INPUTS_FORMAT is unchanged. Fixes One line and exit 1 instead of a traceback: an empty or absent --quantdir, a sample map without a donor column, a per-donor table without a class's _TPM column, a stats --condition no donor has, a JSON input that is not UTF-8, and a qc config without a complete contamination_qc. Symbols are quoted into request URLs. A symbol with a space or a slash made a URL that could not be sent, which was retried and then reported as a network failure. A gzipped --background-fasta is told by its first two bytes, not by its name. A configured transcript the pinned release lacks is named before the gene background is judged. --save-inputs checks that it can write before the run and saves after the report. Saved inputs say where each sequence came from (sequence_sources). No "reads are split" warning for a same-name copy whose sequence is a configured transcript's: the NOTE beside it already said such a record is not counted. extract names a transcript that two groups share; its TPM still goes where 2.1.1 put it. annotate --json keeps its notes, on stderr. The weekly Ensembl check tells a retired archive from an outage. scripts/01_salmon_quant.sbatch writes one output directory per sample map and builds ENA paths for run numbers of 6 to 9 digits. Documentation that said what the code does not do: the Wilcoxon test is SciPy's method="auto"; only the per-cohort and pooled tests carry a floor; a background scan's memory is set by its longest record; several docs/api.md entries. Known issues, for 2.5.0 Records of --background-fasta change the unique k-mer counts but do not enter the compatibility system, so that route can give a more optimistic exit status than the same sequences given with --background-sequences (#14). A gene-background transcript identical to a configured one still counts as a competitor (#15). A temporary redirect off the REST service reads as a retired archive (#12). Corrections to earlier notes The GitHub notes of 2.3.0 said "6 of 36"; it is 8 of 36, as the 2.4.0 changelog says. The changelog's new "Corrections to earlier entries" lists five sentences of the 2.2.0, 2.3.0 and 2.4.0 entries that are wrong, with what is true instead. Provenance Version 2.1.1, cited for the LEPR choroid-plexus analysis, is untouched, as is its Zenodo record 10.5281/zenodo.20738150. extract's transcript-to-group mapping and per-donor TPM summation are unchanged, and its rows are in 2.1.1's order again. The bundled self-test (isoform-dominance selftest) reproduces the reference LEPR result.