Skip to content

PREreview of "When Does a Second Model Help? Cross-Model Review in LLM Verification"

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23132634. Thank you for this paper. You took 30 Korean-language artifacts (code modules, tutorials and presentation scripts, generated by Claude Opus 4.6) with 150 planted errors, and ran 900 review sessions in 10 conditions. The top cross-model reviewer (GPT-5.4) is not significantly different from same-model review in a fresh session, and you note this is not equivalence. At two review calls, one same-model plus one cross-model review matches 56.7% of planted errors against 42.7% for two same-model reviews, but not significantly more than two top-tier cross-model reviews. I liked the audit of session records before the analysis. One baseline run was dropped because most of its findings appear verbatim in the script that wrote the files, and all-session values are still reported. And the paper says plainly what each test does not show. I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents, so maybe my practical side is useful here. Four comments. 1. On the reviewer's task - each reviewer returns a structured list of findings (Section 4.1, Procedure), but the review prompt itself is not quoted. Could you publish it? In my guides I write that a second, fresh agent told to disprove the work ("treat it as wrong until shown otherwise") finds what the first one missed. The two approaches overlap only partly (Jaccard 41.2%, Section 5.2). Could you add one condition where the same-model fresh reviewer gets that adversarial instruction, at the same number of calls? Then we would see how much of the cross-model gain a change of stance alone gives. 2. On the false positives - in Section 5.8, 187 of 1,995 cross-model false positives were classified as model bias, and in 136 of them the reviewer flagged generator-specific features or tools as nonexistent. The classification is keyword-based and not validated. Also, Appendix A counts a finding that correctly identifies a pre-existing (not planted) defect as a false positive too. In my guides I write that a model sounds equally sure when it is right and when it is wrong. Two suggestions. Report the severity label (Section 4.1) separately for true and false positives, so readers can see whether it helps a human decide which findings to check first. I expect it may not help much. And check a random sample of false positives by hand, to see how many are real pre-existing defects. 3. On who wrote the answer key - according to Appendix A, the errors were injected by the generator model itself (Claude Opus 4.6), and the ground-truth records were written at injection time. The same model did the same-model reviews (Section 3.2). In Section 5.4, CCR has F1 40.7% against 37.2% for XMR-GPT on code, and your conjectured explanation is a shared understanding of programming patterns. I'd add a second reading. The advantage may partly come from the same model having written the planted errors. In my own experiment (36 runs, DOI 10.5281/zenodo.22759217) the agent wrote its own tests, and the gate returned 4,086 tests and 0 failures. The agent wrote the criterion and then met it itself, so all-green told me nothing. A small set of artifacts with errors injected by a different model or by a person would separate the two explanations. 4. On cost and open data - Section 2.5 matches review calls, not compute. In my measurement of 722 agent sessions (DOI 10.5281/zenodo.22759216), more than 85% of modeled cost was context work (cache reads plus cache writes). So I'd add input and output tokens per review session for each condition. Information restriction also changes the input, so then Section 6 can be read as cost per found error, not per call. The Limitations also say CCR runs differ widely (56, 37 and 29 matched errors) and the headline pairing is among the higher run combinations. In my own experiment I also chose the comparison point after the runs, and said so in the report. My run data is public on Hugging Face (DOI 10.57967/hf/10366). I'd suggest posting the per-session records publicly with a DOI instead of on request, so the set-level numbers can be recomputed by anyone. Thank you for a careful paper. Competing interests Yes: the text cites the author's own technical reports and dataset (DOI 10.5281/zenodo.22759216, 10.5281/zenodo.22759217, 10.57967/hf/10366); no connection to the preprint's author. Use of Artificial Intelligence (AI) The author declares that they did not use generative AI to come up with new ideas for their review.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#computer vision Review Mar 2008

Agile methods in European embedded software development organisations: a survey on the actual use and usefulness of Extreme Programming and Scrum

The results show that the embedded industry has been able to apply agile methods in its development processes and that the appreciation of the agile methods and their individual practices appears to increase once adopted and applied in practice.

O. Salo, P. Abrahamsson · 238 citations · ⚡9
#computer vision Open access Jul 2017

What happens when software developers are (un)happy

Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 236 citations · ⚡13

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.