Skip to content

brunetch123/K2-Annotator: K2 Annotator v3.1.0

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Release produced from an independent scientific and software review of v3.0.21 against the manuscripts the pipeline implements (Koelmel et al. 2022, Brunet et al. NYCSS manuscript + SI, Kwiecien et al. 2015). Finding IDs (S-* scientific, D-* design) refer to that review; every finding is pinned by a test in the new tests/ package (python -m pytest). No Level-2 threshold or window was changed. Several fixes make the code compute the written criteria correctly and therefore change scores; see "Migration". Fixed — scoring (changes results) RHRMF now subtracts the electron mass (S-RHRMF-1). rhrmf.py compared measured ion m/z against neutral sub-formula masses. Kwiecien 2015 (SI) subtracts one electron mass before matching. The omission shifted the ±10 ppm window by 6 ppm at m/z 91 and 11 ppm at m/z 51, rejecting perfectly measured light fragments and making the effective tolerance asymmetric (toluene at −5 ppm mass error scored 0, at +5 ppm scored 100). The window is now symmetric about the measured cation m/z. RHRMF values rise for spectra rich in fragments below ~m/z 150; the "unexplained biphenyl zeros" noted in the v3.0.20 entry are a symptom of this defect. Isotopically labelled formulas work in RHRMF (S-RHRMF-2). D/2H, 13C and other isotope symbols emitted by molmass were silently dropped, so every labelled surrogate scored RHRMF 0 against a low-resolution entry. Atom masses now come from molmass' isotope table (also fixing a ~0.8 ppm bias from the old 5-decimal table, S-DOC-1); unknown atoms trigger a warning at library load instead of silence. Exact-mass library detector decides by stored precision (S-HR-1). An entry is HR when its metadata says so (resolution/hires = high) or when ≥2 peaks carry ≥3 decimal digits with ≥1 in the top-3 by intensity. The v3.0.21 mass-defect rule classified exact-mass benzene, chlorobenzene and dichlorobiphenyl entries as low-resolution and discarded real C27–C28 alkyl fragments as z=2 artefacts. The old rule remains as is_library_high_res_v3021() for audit. Surrogate recoveries were IS-normalised twice (S-SUR-1). With IS normalisation on, SurrogateAnalyzer multiplied already-normalised abundances by the factor again (factor² → a true 100 % recovery reported as 200 %/400 %). The surrogate matcher also now evaluates the 1.5 % RI rule relative to the feature RI like the screening engine (S-SUR-2). Unit-mass binning rounds half up (76.5 → 77 and 77.5 → 78; previously banker's rounding gave 76 and 78) (S-DOT-2). Fixed — correctness --grouping is live (D-1). The GUI's per-sample Blank/Sample/Reference table was written to JSON and forwarded, but cli.py never used it; BFF fell back to the blank-identifier substring. Reference samples are derived from the grouping as well. Zero-of-anything is an error (D-2, D-5, S-BFF-1, S-RI-2). No quantification columns, no features, no sample columns, no blank columns (unless --allow-no-blanks), unknown names in the grouping, a missing calibration file, or MZmine data without a calibration file now stop the run with a clear message instead of an empty "successful" result. The final blank/sample classification is printed. One tolerant MSP reader (src/msp_reader.py) for every consumer (D-3). Handles UTF-8 BOM, CRLF, any key case, Retention_index:/NUM_PEAKS: spellings, NIST 41 59; 43 999; and (41 59) peak lines, records without blank-line separators, and Num Peaks mismatches (warned, never truncated). Previously a NIST export could load zero compounds silently. MZmine quant files with Peak height or datafile:X.mzML:area columns are accepted; sample names are normalised (extension / suffix stripped) so GUI file stems, --reference-samples and column headers agree. RI calibration loader validates its input (D-4): header-less tables no longer lose their first alkane; comma/semicolon/space delimiters accepted; duplicate carbon numbers and non-monotonic RTs are rejected with explicit messages. IS normalisation is idempotent and validated (D-10): factors are applied to the parsed (raw_abundances) values, string IS values fail with a message, the IS MSP goes through the shared reader, auto-detected IS areas are reported in IS_Area_* columns. Reports (D-9): Instrument/Comments columns read the template keys instrument/comments; reference samples are excluded from feature-summary statistics (they were excluded from the BFF but not from the summaries); Mean_Detected_Abundance added; hazard lookups happen once per CAS; mirror plots are rendered only for the PDF; one bad page no longer loses the whole PDF (c.save() in finally); temp_assets/ is removed unless --keep-assets; surrogate recovery of exactly 0 renders as 0. External APIs (D-8): one requests.Session with timeouts, retries and a PubChem rate limit; per-CAS on-disk cache (~/.k2/cache/); the guessed EpaClient endpoint (which returned an all-zero matrix and suppressed the PubChem fallback whenever an API key was present) is removed; DTXSIDs come from ctxpy; pubchempy dependency dropped. GUI / pipeline (D-7): the Thread Count setting is honoured (import serialisation is now an explicit checkbox); the PyInstaller build runs the pipeline in-process instead of re-launching the GUI; max_lib_peaks, ri_extrapolation and allow_no_blanks are exposed in the Analysis screen; project/preset state no longer leaks between projects; the EPA API key travels via K2_EPA_API_KEY and is stored only in ~/.k2/credentials.json, never in .K2 files or command echoes; Tk widgets are updated on the main thread; Cancel kills the whole process tree; results/ is never overwritten (_2, _3, …); --name is sanitised and used; dead code (main.py, InternalStandardScreen, msdial_parser.py) removed. Added RI_Extrapolated column in every table and a load-time warning: RIs for features outside the alkane range are flagged (S-RI-1). --ri-extrapolation {spline,linear} selects how they are derived (default spline = v3.0.x behaviour; linear = van den Dool from the terminal alkane pair). RICalibrator.rt_range / is_in_range(). run_manifest.json in every results folder (D-6): K2 and package versions, SHA-256 of every input, resolved options (including max_lib_peaks), sample classification, library statistics, calibration range and feature/match counts. pipeline_log.txt and the untruncated mzmine_log.txt are saved alongside. HR_RevDot / HR_FwdDot columns; RHRMF reads N/A for exact-mass entries instead of the sentinel 100.0 (S-HR-2). CLI exit codes: 0 matches found, 2 completed with no matches (summaries and manifest still written), 1 error. --csv-only/--pdf-only are mutually exclusive; --version. tests/: 99 pytest tests with a formula-derived synthetic dataset (tests/synth/make_dataset.py) covering every criterion, edge cases and the end-to-end CLI. templates/ri_calibration_template.txt ships so the template data runs out of the box. docs/SCORING_METHODS.md: the exact computations (weights, trimming, detector rule, RHRMF details, extrapolation) in SI-ready form. Documentation Dot-product weighting (m/z¹ · I^0.5, squared cosine × 1000), the top-20 library trim and the HR-aware re-scoring step are now stated explicitly; they were undocumented (S-DOT-1, S-TRIM-1). The adjusted BFF mode is labelled as outside the Koelmel framework. The IS factor is documented as max(IS)/IS (the guide said median). Migration Re-run analyses whose RHRMF values, HR classification or surrogate recoveries matter: RHRMF scores rise for light-fragment spectra, some candidates rejected by v3.0.x now pass, exact-mass aromatic/halogenated entries take the HR path, and IS-normalised recoveries are lower by the IS factor. MZmine runs now require --ri-cal; runs with no blank column require --allow-no-blanks. epa_client.py, msdial_parser.py and main.py are gone; pubchempy is no longer required. Full test suite: python -m pytest (99 tests). Methods as implemented: docs/SCORING_METHODS.md.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54

Related blog posts

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

Microsoft Research Blog Jul 30, 2026

Echoverse: Deep, evolving environments for computer-use agents

Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve. The post Echoverse: Deep, evolving environments for computer-use agents appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.