Skip to content

The Reassessment That Did Not Travel: What the 2020 Arcade Result About Exploration Bonuses Establishes, the Six Conditions That Made It Informative, and Which of Them the Language-Model Revival Drops

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

(c) 2026 Pranay Mahendrakar. Licensed under CC BY 4.0. Exploration bonuses have returned. Between 2025 and 2026 at least six frameworks added an intrinsic novelty or uncertainty term to reinforcement learning with verifiable rewards for language models, each naming a classical antecedent - prediction error, pseudo-counts, epistemic uncertainty - and each reporting gains. The classical literature those antecedents come from also contains a controlled reassessment. In work published at ICLR 2020, a study held the learning algorithm fixed, tuned every bonus, and compared against plain undirected exploration across the full Atari suite; it reported that bonuses beat the simple scheme on one celebrated game, showed no visible difference from it on the rest of the designated hard-exploration set, and never beat it on games where exploration is not the bottleneck. This paper states what that reassessment establishes and, at comparable length, what it does not; extracts the six design conditions that made it informative; and audits the language-model revival against them. Sixteen papers in the revival were checked mechanically for a citation to it, and none contains one. Four of the six conditions are met by at least one paper in the revival and a fifth in part. The condition the reassessment was built to test - an evaluation arm where exploration is not the bottleneck - is met by none of them in the form it requires, although the pattern that condition exists to detect is already visible in one revival paper's own published table. This paper reports no experiments. It names two failed direct imports that the revival itself reports and does not read as evidence about transfer, states where the analogy breaks on the substrate rather than on the evidence, and specifies the comparison a 2026 survey independently asks for without knowing it has been run once already. The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv record before inclusion, and every quantitative claim was read back against the cited source's own abstract or, where a claim is drawn from a paper's body or a published table, against the located passage. The absence claims in Section 7 were produced by fetching each named paper's full text and searching it mechanically; the procedure is stated in Section 2 so that it can be repeated. The author is responsible for the final text and for all claims made in it.

View source

Similar papers

#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#machine learning Review Open access Jun 2014

Why Early-Stage Software Startups Fail: A Behavioral Framework

This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.

Carmine Giardino, Xiaofeng Wang, P. Abrahamsson · 175 citations · ⚡19
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#machine learning Review Open access May 2016

Key Challenges in Software Startups Across Life Cycle Stages

It is found that what perceived as biggest challenges by software startups do vary across different life cycle stages, even though its significance decreases when the learning focuses of the startups move from problem to solution and their products mature.

Xiaofeng Wang, Henry Edison, Sohaib Shahid Bajwa et al. · 62 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.