Sep 2026· Zenodo (CERN European Organization for Nuclear Research)
Abstract
The reproduction workflow and the archived intermediate results for the paper. Everything needed to rebuild the manuscript's tables is here: the code is a snapshot of the trimmed reproduction repository, and the bundles hold the computed intermediates, so a reader does not have to re-derive them. Reproducing from the raw corpora instead needs roughly 200 GB of intermediates and several GPU-days. doc2lora-embedding-code- .tar.gz — the workflow itself: a Snakemake pipeline trimmed to the dependency closure of what the manuscript reads (24 rule files; snakemake -n paper_assets resolves 330 jobs from the raw corpora). Includes REPRODUCE.md, which maps every figure, table and quoted number to the rule that produces it, and data/ARTIFACTS.tsv, a SHA-256 per archived file. doc2lora-results.tar.zst (1227 files) — everything the manuscript's tables and figures are computed from that needs a GPU, the licensed APS text, or an LLM judge to produce: the paired per-unit score pools behind the bootstrap confidence intervals, the invertible adapter weights, every method's raw cluster labels and the judge verdicts, the pair-axis decodes (Doc2LoRA and ICAE), the PACS node set, and the manuscript's own tables for comparison. With this bundle alone, snakemake paper_assets --rerun-triggers mtime rebuilds every table and figure the paper reads, running only CPU rules. doc2lora-s2and.tar.zst (50 files) — the gene and text embeddings for the five author-name-disambiguation benchmarks (zbMATH, QIAN, ArnetMiner, PubMed, KISTI), so the disambiguation rows can be re-scored from the vectors up. Unpack the code snapshot, then fetch and verify the bundles against data/ARTIFACTS.tsv (SHA-256 per file): tar xzf doc2lora-embedding-code- .tar.gz && cd doc2lora-embedding- cp workflow/config.template.yaml workflow/config.yaml # edit the paths python scripts/fetch_artifacts.py results --record snakemake paper_assets -j4 --rerun-triggers mtime # rebuilds every table, CPU only The 644k-paper APS gene matrices (~137 GB) are not included: they exceed a Zenodo record and are a deterministic function of the APS corpus plus the published hypernetwork checkpoints (snakemake all_embeddings). The APS corpus itself is licensed and is not redistributed here.
Some claim that especially in the field of agile software development the research lags years behind of the practice. In this paper, we characterize the status and main challenges for research on agile software development, and propose a preliminary roadmap, focusing on providing more empirical research, primarily on e...
Torgeir Dingsøyr, T. Dybå, P. Abrahamsson· Agile Conference· 92 citations· ⚡7
The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· International Conference on...· 84 citations· ⚡6
This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.
Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al.· Journal of Systems and Softw...· 78 citations· ⚡6
The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.
Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al.· arXiv.org· 62 citations· ⚡3
Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...
A. Marchenko, P. Abrahamsson· Agile Conference· 59 citations· ⚡11
A comprehensive overview of how enhanced sampling methods are reshaping the field, with a particular focus on the data-driven construction of collective variables, is provided.
Kai Zhu, Enrico Trizio, Jintu Zhang et al.· Chemical Reviews· 58 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.