Skip to content
#small language model Open access

PREreview of "When Does a Second Model Help? Cross-Model Review in LLM Verification"

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23132634. Thank you for this paper. You took 30 Korean-language artifacts (code modules, tutorials and presentation scripts, generated by Claude Opus 4.6) with 150 planted errors, and ran 900 review sessions in 10 conditions. The top cross-model reviewer (GPT-5.4) is not significantly different from same-model review in a fresh session, and you note this is not equivalence. At two review calls, one same-model plus one cross-model review matches 56.7% of planted errors against 42.7% for two same-model reviews, but not significantly more than two top-tier cross-model reviews. I liked the audit of session records before the analysis. One baseline run was dropped because most of its findings appear verbatim in the script that wrote the files, and all-session values are still reported. And the paper says plainly what each test does not show. I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents, so maybe my practical side is useful here. Four comments. 1. On the reviewer's task - each reviewer returns a structured list of findings (Section 4.1, Procedure), but the review prompt itself is not quoted. Could you publish it? In my guides I write that a second, fresh agent told to disprove the work ("treat it as wrong until shown otherwise") finds what the first one missed. The two approaches overlap only partly (Jaccard 41.2%, Section 5.2). Could you add one condition where the same-model fresh reviewer gets that adversarial instruction, at the same number of calls? Then we would see how much of the cross-model gain a change of stance alone gives. 2. On the false positives - in Section 5.8, 187 of 1,995 cross-model false positives were classified as model bias, and in 136 of them the reviewer flagged generator-specific features or tools as nonexistent. The classification is keyword-based and not validated. Also, Appendix A counts a finding that correctly identifies a pre-existing (not planted) defect as a false positive too. In my guides I write that a model sounds equally sure when it is right and when it is wrong. Two suggestions. Report the severity label (Section 4.1) separately for true and false positives, so readers can see whether it helps a human decide which findings to check first. I expect it may not help much. And check a random sample of false positives by hand, to see how many are real pre-existing defects. 3. On who wrote the answer key - according to Appendix A, the errors were injected by the generator model itself (Claude Opus 4.6), and the ground-truth records were written at injection time. The same model did the same-model reviews (Section 3.2). In Section 5.4, CCR has F1 40.7% against 37.2% for XMR-GPT on code, and your conjectured explanation is a shared understanding of programming patterns. I'd add a second reading. The advantage may partly come from the same model having written the planted errors. In my own experiment (36 runs, DOI 10.5281/zenodo.22759217) the agent wrote its own tests, and the gate returned 4,086 tests and 0 failures. The agent wrote the criterion and then met it itself, so all-green told me nothing. A small set of artifacts with errors injected by a different model or by a person would separate the two explanations. 4. On cost and open data - Section 2.5 matches review calls, not compute. In my measurement of 722 agent sessions (DOI 10.5281/zenodo.22759216), more than 85% of modeled cost was context work (cache reads plus cache writes). So I'd add input and output tokens per review session for each condition. Information restriction also changes the input, so then Section 6 can be read as cost per found error, not per call. The Limitations also say CCR runs differ widely (56, 37 and 29 matched errors) and the headline pairing is among the higher run combinations. In my own experiment I also chose the comparison point after the runs, and said so in the report. My run data is public on Hugging Face (DOI 10.57967/hf/10366). I'd suggest posting the per-session records publicly with a DOI instead of on request, so the set-level numbers can be recomputed by anyone. Thank you for a careful paper. Competing interests Yes: the text cites the author's own technical reports and dataset (DOI 10.5281/zenodo.22759216, 10.5281/zenodo.22759217, 10.57967/hf/10366); no connection to the preprint's author. Use of Artificial Intelligence (AI) The author declares that they did not use generative AI to come up with new ideas for their review.

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.