Skip to content
#software testing Open access

Analysis code and supporting materials for 'Same Practice, Opposite Ranking: Disciplinary Inversion in Research Integrity Norms across Cameroonian Universities'

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

Materials supporting “Same Practice, Opposite Ranking: Disciplinary Inversion in Research Integrity Norms across Cameroonian Universities”. The article stands on its own: everything a reader needs to follow and judge its argument is in it. This deposit is not a continuation of the article. It exists so that anyone who wants to check the analyses, rather than read them, can do so, and so that the route to the reported result is on record. Contents analysis.py — One script. Reads the two source files and reproduces every value reported in Section 4 of the article: the 34 pre-specified tests with their Benjamini-Hochberg adjustments, the four pre-specified logistic models, the four sensitivity analyses, the two sensitivity tables, and the reported percentages. outputs_appendix_a_inventory.csv — The 34 tests: family, n, statistic, degrees of freedom, raw and adjusted p, Cramér’s V, smallest expected cell count. outputs_appendix_b_family_sensitivity.csv — For each hypothesis family, which tests survive adjustment and over what range of family sizes that set is unchanged. outputs_table_4_5_combinations.csv — The six combinations behind Section 4.5 of the article. outputs_models.csv — The eight logistic models. outputs_prevalences.csv — The percentages of Tables 4 and 5 and of Sections 4.2, 4.4 and 4.5, each with its numerator, its denominator, and the base it is computed on. Supplementary_material.docx — The two survey instruments, the pre-specified analysis plan, its dated addendum, the model tables and the inventory of tests. This is the file submitted to the journal with the manuscript. errors_and_corrections.md — The errors found in the original analyses before submission, and how each was corrected. run_log.txt — The console output of a complete run on the source data, ending with the check against the reported values. test_analysis.py — Checks that the analysis rules behave as Section 3.6 states, on ten rows built for the purpose. No data needed. LICENSE.txt — MIT for the code, CC BY 4.0 for the documents and outputs. Zenodo stores files without folders, so the five outputs carry an outputs_ prefix here. The script writes them under their bare names into the working directory. One reported figure is reproduced by none of them. The response rate of Table 2, 34.5%, divides the 207 staff respondents by the number of academics convened at the sessions, and that denominator is recorded in neither source file. The data are not here The two source files hold individual survey responses. Neither questionnaire collected direct identifiers. The staff questionnaire records university and discipline, the doctoral questionnaire university and establishment, the latter serving only to assign respondents to a broad domain. The university enters no analysis and is reported for no respondent; no result is reported for, or attributed to, any university or establishment. Respondents were told that their answers would be analysed in aggregate and the findings alone disseminated. No consent to redistribution was sought, so the files are not deposited; requests are considered for the purpose of verifying the analyses reported in the article, under an agreement excluding redistribution and any other use. The consequence is worth stating plainly. This deposit lets anyone verify the analysis. It does not let anyone verify the collection. No one outside the study can check that the two files correspond to what respondents actually ticked, and no deposit can supply that: the chain of verification in empirical work ends on a human attestation. Here it ends in December 2022, during a session break, when someone marked a paper questionnaire. What inspection alone does give is considerable, and it is what the principal error turned on: which variables each test crosses, which coding rule it applies, and how each family is adjusted are all legible in the code without executing it. Running it pip install pandas numpy scipy statsmodels openpyxl python analysis.py Sondage.xlsx DOCTORANTS_2024.xlsx Expected inputs: Sondage.xlsx, sheet Exported, 207 academic staff; and DOCTORANTS_2024.xlsx, sheet Données, 317 doctoral students. The script prints the inventory and the reported percentages, lists what survives adjustment, checks six values against those reported, and ends with All checks passed. when the run agrees with the article. If it ever stops saying so, something has moved. run_log.txt is the output of that run, so the 34 tests, their adjustments and the reported percentages can be read against Section 4 of the article without executing anything and without access to the data. The models and the sensitivity analyses are in outputs_models.csv. test_analysis.py runs without the data at all: python test_analysis.py It calls the functions of analysis.py on ten rows built by hand, and checks six behaviours of the analysis rules: that ‘no opinion’ answers are excluded rather than recoded, and that this changes the base of the test; that blank responses are dropped variable by variable rather than listwise; that Yates’ correction applies to 2x2 tables; that Fisher’s exact test takes over below an expected count of five; that degrees of freedom follow the shape of the table; and that Cramér’s V comes from the uncorrected Pearson statistic and cannot be recovered from the Yates statistic printed beside it. It also checks that the Benjamini-Hochberg step-up divides by rank, enforces monotonicity and preserves the input order, and, on the family of Section 4.5, that four members give .0466 where five give .0582, so that the size of the family decides whether that association survives adjustment. Those ten rows are not a sample of the survey. Analysis rules The rules stated in Section 3.6 of the article are written out at the head of analysis.py as R1 to R10 and cited where each applies. In brief: Pearson’s chi-square, with Yates’ correction on 2x2 tables and Fisher’s exact test where an expected count falls below five; Cramér’s V computed throughout from the uncorrected Pearson statistic, so that values stay comparable across tables of different dimensions; Benjamini-Hochberg by hypothesis family at 5%, step-up with monotonicity; ‘no opinion’ retained as a distinct category in the prevalence table of Section 4.1 and excluded from binary association tests; blank responses treated as missing variable by variable, never listwise. Two consequences of the rule on Cramér’s V are easy to mistake for errors. Cramér’s V cannot be recovered from a chi-square printed with Yates’ correction, and for the one association tested by Fisher’s exact test it derives from a statistic not printed at all. Provenance The analyses were produced by Claude (Anthropic), which also drafted the analysis plan. Every value was recomputed twice from the source files: by the author using other software, and in a separate session by the same assistant, working from the data rather than from the author’s summaries. The second recomputation is not independent of the assistant that produced the original analyses; the author’s is. Both identified errors, and the author’s corrected several in the assistant’s. This script is the consolidated result. errors_and_corrections.md gives the full account.

View source

Similar papers

#computer vision Review Sep 2017

Agile Software Development Methods: Review and Analysis

This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.

P. Abrahamsson, O. Salo, Jussi Ronkainen et al. · 727 citations · ⚡54
#computer vision Jun 2008

The impact of agile practices on communication in software development

The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.

M. Pikkarainen, Jukka Haikara, O. Salo et al. · 401 citations · ⚡48
#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#computer vision Review Mar 2008

Agile methods in European embedded software development organisations: a survey on the actual use and usefulness of Extreme Programming and Scrum

The results show that the embedded industry has been able to apply agile methods in its development processes and that the appreciation of the agile methods and their individual practices appears to increase once adopted and applied in practice.

O. Salo, P. Abrahamsson · 238 citations · ⚡9
#computer vision Open access Jul 2017

What happens when software developers are (un)happy

Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.

D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al. · 236 citations · ⚡13
#computer vision Open access Oct 2004

Mobile-D: an agile approach for mobile application development

The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.

P. Abrahamsson, Antti Hanhineva, H. Hulkko et al. · 225 citations · ⚡18

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.