CorVer (Corpus Verify) is a lightweight training-time approach to process supervision that derives sentence-level rewards from corpus co-occurrence statistics and outperforms the unmodified models in all standard factual-QA settings.
Abstract
Process supervision during reinforcement learning (RL) post-training matters in factual question answering (QA) because responses receiving positive response-level rewards can still contain sentence-level factual errors. Unlike math and code, factual QA lacks inexpensive programmatic checks, making reward computation a bottleneck. Existing methods rely on neural verifiers to score individual sentences, requiring extensive model inference as RL repeats these checks across many sampled responses at every update. We therefore introduce CorVer (Corpus Verify), a lightweight training-time approach to process supervision that derives sentence-level rewards from corpus co-occurrence statistics. A 0.5B extractor identifies subject-object pairs, indexed corpus queries supply their co-occurrence counts, and the resulting sentence rewards are assigned to the corresponding tokens for RL. On Qwen3-4B and Qwen3-8B, CorVer reduces mean complete training time by 5.5-10.4 times relative to the four factuality-RL baselines. Across the four models evaluated against these baselines, CorVer achieves the highest accuracy in 17 of 20 model-benchmark settings. CorVer outperforms the unmodified models in all 30 standard factual-QA settings (six models from three families across five benchmarks) and remains effective on two additional multi-hop QA datasets.
This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.
P. Abrahamsson, O. Salo, Jussi Ronkainen et al.· arXiv.org· 727 citations· ⚡54
Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· Journal of Systems and Softw...· 236 citations· ⚡13
The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.
P. Abrahamsson, Antti Hanhineva, H. Hulkko et al.· Conference on Object-Oriente...· 225 citations· ⚡18
A study with 42 participants investigates the relationship between the affective states, creativity, and analytical problem-solving skills of software developers and offers support for the claim that happy developers are indeed better problem solvers in terms of their analytical abilities.
D. Graziotin, Xiaofeng Wang, P. Abrahamsson· PeerJ· 216 citations· ⚡13
GAOKAO-Bench is introduced, an intuitive benchmark that employs questions from the Chinese GAOKAO examination as test samples, including both subjective and objective questions that contribute a robust evaluation benchmark for future large language models and offers valuable insights into the advantages and limitations...
Xiaotian Zhang, Chun-yan Li, Yi Zong et al.· arXiv.org· 216 citations· ⚡17
This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.
Carmine Giardino, Xiaofeng Wang, P. Abrahamsson· International Conference on...· 175 citations· ⚡19
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.