Skip to content

NP-SEE v6.1: Genomic Structural Analysis of Filoviruses and Emerging Pathogens -- Ebola Bundibugyo 2026 Epidemic Strain, Marburg Angola, CCHF, VZV, PRV

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)
Viral Infections and Outbreaks Research

Abstract

Part one (new in this version, pre-registered). Forty-nine complete filovirus genomes from nine species were compared against order-2 and order-5 Markov models fitted to each 16,384-base window, with three surrogate replicates per window. Across the 19 genomes frozen as the primary set, all four declared measures depart from the model in the same direction (median differences +0.0657, +0.0172, +0.0520, +0.1707; Holm-corrected p = 1.5e-05). An independent replication on 30 genomes acquired afterwards, with the success rule fixed in writing before acquisition, confirms all four (Holm-corrected p = 7.5e-09). Replication magnitudes are systematically smaller than the primary ones, so the primary estimates are reported as upper bounds. A direct noise measurement, from the spread of the surrogate replicates, ranks the four measures by reproducibility. Part two (pre-registered). Whether the G-quadruplex motifs of the Bundibugyo lineage remained identical between the 2007 and the 2026 outbreaks. Under a strict motif definition the 2007 reference genome contains two motifs; one is identical in all 21 targets, against 139 of 208 control segments identical in all targets. With two motifs the comparison has no power and the outcome is inconclusive. A gap in the design is declared: no minimum number of motifs had been fixed for the primary comparison, only for a secondary one. Withdrawals. The claim made in version 4, that genomic structure was conserved across three outbreaks and nineteen years, is withdrawn, together with the section listing candidate structural targets that rested on it. The underlying density result (GQ-like counts 35.0, 36.0 and 37.6 across the three corpora) stands and is retained; what does not follow is the step from a stable count to conserved structures, and from there to targets. A negative control added in this version also withdraws the MICRA label. Shuffled genomes, which retain length and base composition but have no structure, obtain the same label as the real genomes: nine shuffles out of ten reach the strongest category for the two large genomes, and no pathogen separates from its own shuffles (empirical p from 0.097 to 0.742). The label tracks the number of windows examined rather than structure. Analyses from versions 1 to 4 on CCHF IbAr10200, VZV, PRV and the Marburg clades are retained as descriptive measurements with their validation controls, and are marked as not redone under pre-registration. Limitations. Computational predictions without experimental validation; no clinical conclusions and no proposed therapeutic targets; the source code is not available for external review; the document has not undergone peer review. Chapter 10 lists every substantive error corrected from version 1 onwards, with the version that introduced it and the version that corrected it. Author: Luca Lupinacci, independent researcher, Cosenza, Italy. The author is not a biologist, a geneticist or a computer scientist by training; the software and this report were produced with the assistance of artificial-intelligence systems. Cite with the concept DOI 10.5281/zenodo.22817161, which always resolves to the latest version. ---------------------------------------------------------------------- IMPORTANT: this version supersedes all previous ones. Readers who downloaded any earlier version are encouraged to use this one. CHANGELOG Version 5 - 2 Oct 2026 New pre-registered part on 49 genomes with an independent replication; the MICRA label withdrawn after a negative control; the claim about conservation of structure withdrawn; the section on targets removed; laboratory constructs excluded; reliability measurement added; window overlap declared; errors chapter added; numbering and DOI made consistent. Version 4 - 24 Sep 2026 - 10.5281/zenodo.22936607 MICRA recomputed at 1000 permutations for all 5 pathogens; cross-validation with QGRS Mapper and G4Hunter on BDBV and Marburg; positive controls on c-MYC and human telomere; Hamming distance analysis on the 2007 corpus; complete GC/N/GQ table for the 8 Zaire 2025 genomes; M32 moved to an appendix; genomic files re-verified. Version 3.1 A document revision, not published as a Zenodo version: the "pan-filoviral conserved" region renamed "candidate"; target language revised after independent review; exploratory thresholds declared as pre-defined. Version 3 - 20 Sep 2026 - 10.5281/zenodo.22861938 Substantive corrections: 3 genomic files not matching their accessions replaced; FJ217162 removed from statistical comparisons; MICRA recomputed at 1000 permutations; GQ-like positive control added; M32 removed; "GQ" renamed "GQ-like"; 95% CI for the 2026 corpus corrected to bootstrap percentile. Version 2 - 19 Sep 2026 - 10.5281/zenodo.22846534 Corpus and analyses extended. Version 1.0 - 17 Sep 2026 - 10.5281/zenodo.22817162 First publication. NOTE ON NUMBERING: version 3.1 was a document revision and was never published as a separate Zenodo version; it has no DOI of its own. From version 5 there is a single numbering scheme, no sub-versions, and the document cites only the concept DOI. FILES: the report in English and Italian, the pre-registration with its dated amendments, the verified genome files with checksums, the per-genome measurements, the motif table, the quality-control tables, the negative-control output, and a manifest listing every file with its SHA-256. The source code is not included.

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#artificial intelligence Open access May 2023

Evaluating the Performance of Large Language Models on GAOKAO Benchmark

GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations...

Xiaotian Zhang, Chun-yan Li, Yi Zong et al. · 216 citations · ⚡17
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#artificial intelligence Open access Jul 2024

Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval

This work investigates the possibilities of using LLMs in a resume screening setting via a document retrieval framework that simulates job candidate selection and finds that the MTEs are biased, significantly favoring White-associated names in 85% of cases and female-associated names in only 11.1% of cases.

Kyra Wilson, Aylin Caliskan · 131 citations · ⚡8
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.