Skip to content
#software testing Open access

Rubrica v0.1.5: software description and validation report

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

This report documents Rubrica v0.1.5, a local application for grading scanned, handwritten exams against an instructor's rubric. Under the default local-Ollama configuration, cover-page extraction runs on the instructor's machine. Stored names, SIDs, and page 0 are excluded from external grading requests; those requests contain an anonymous identifier, answer-page images, and rubric text. Rubrica does not inspect answer pages or rubrics for incidental identifiers, so instructors should review and remove that information before external grading. Version 0.1.5 is the current pilot build (September 2026), a signed and notarized standalone macOS application for Apple Silicon Macs on macOS 12 or later (Tauri shell around the Flask backend, React interface). Its grading and scoring logic is unchanged from version 0.1.2, the first pilot build distributed to outside testers in late August 2026. The validation evidence below is carried forward from the previous version of this record; the studies were run on development builds between April and July 2026, before the 0.1.2 pilot build, while parts of the grading prompt and its safeguards were still changing. This report contains a software description, a validation note, architecture and grading-pipeline diagrams, a README, and citation metadata. It does not contain source code or an executable. The deposited documents and diagrams are released under CC BY 4.0. That license does not cover the absent AGPL core source or proprietary pilot application. Validation evidence. An author-run, matched-input cross-model comparison (April 2026; 30 exams, 840 scored items; Claude Sonnet 4.6 vs. o4-mini) reported ICC(3,1) 0.964, quadratic weighted kappa 0.883 on letter grades, 96.2% within-one-point agreement, and mean absolute error of 0.15 points per question. The observed production-minus-reference mean difference was +0.64 points per exam; its approximate 95% confidence interval spans zero. The comparison estimates cross-model concordance and serves as a structural check, not a third-party audit or human-accuracy study. A July 2026 comparison against human-assigned reference scores covered one 25-exam macroeconomics midterm from another instructor. On the 46.5-point decomposed section, mean absolute difference was 1.26 points per exam, the observed mean difference was -0.28 points, and 84% of item scores matched exactly. The separate 25-point holistic essay had mean absolute difference 3.88 points. These sample-bounded results do not estimate a general accuracy rate or isolate the effect of rubric decomposition. A June 2026 test-retest study selected 10 final-exam submissions by deterministic stratification across version, batch, and score range and graded each three times. All 10 retained the same letter grade; mean within-exam total-score standard deviation was 0.91 points. A separate earlier 30-exam, three-run capture observed 1.03 points; because it used a different corpus and rubric, it is context rather than a formal comparator. The project page at travisfraser.com/rubrica is maintained together with this report. Version 0.1.5 record (2026-09-14). Updated for the 0.1.5 build: the packaged app runs on macOS 12 or later; update checks and first-time sign-in now verify TLS certificates against a bundled certificate store (builds 0.1.3 and 0.1.4 could not complete those connections), and each release is tested for this inside the shipped application before signing; review cards show status as a color wash with the status in words. It also corrects the comparison-model label: the author-run comparison can use Google Gemini 3.1 Pro (the in-app default), Anthropic Claude Opus 4.6, or OpenAI GPT-4o, and the April 2026 study used OpenAI o4-mini (the previous version listed Opus 4.7, which Rubrica uses only to re-check escalated multiple-choice answers). Other changes: in-app comparison metrics no longer withhold QWK and ICC below 30 comparisons, one heading in the rubric auto-enhance prompt was reworded, and the packaged app embeds the python.org build of Python 3.14. Grading and scoring logic and the privacy boundary are unchanged. Earlier revision (2026-09-05). Reclassified the documentation-only deposit as a report; narrowed the privacy statement; replaced determinism with measured test-retest stability; corrected the cross-model method, sample-size guidance, and human-reference wording; and documented the third validation study.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#computer vision Review Mar 2008

Agile methods in European embedded software development organisations: a survey on the actual use and usefulness of Extreme Programming and Scrum

The results show that the embedded industry has been able to apply agile methods in its development processes and that the appreciation of the agile methods and their individual practices appears to increase once adopted and applied in practice.

O. Salo, P. Abrahamsson · 238 citations · ⚡9
#computer vision Open access Jul 2017

What happens when software developers are (un)happy

Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 236 citations · ⚡13
#computer vision Open access Oct 2004

Mobile-D: an agile approach for mobile application development

The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.

P. Abrahamsson, Antti Hanhineva, H. Hulkko et al. · 225 citations · ⚡18

Related blog posts

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

MIT News · Artificial Intelligence Aug 17, 2026

Q&A: Rethinking how innovation happens

In his latest book, Professor Eugene Fitzgerald examines the forces that turn breakthroughs into value — and why innovation resists simple formulas.

Microsoft Research Blog Aug 12, 2026

MindTopo reveals VLMs’ spatial reasoning abilities

A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs’ spatial reasoning abilities appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.