Skip to content
#explainable ai Open access

What a Crossref record shows when a paper is retracted: one month of retraction deposits (v3)

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research) · 2 references
Academic integrity and plagiarism

Abstract

# What a Crossref record shows when a paper is retracted **A field-level audit of one month of retraction deposits (v3)** Author: Trafalgar Law (independent). Data pulled from the public Crossref REST API on 2026-09-23. Sample and reproduction query included. This is a metadata audit, not an accusation against any publisher or author. ## Why On 2026-09-19 the record for a fisetin/renal-fibrosis paper in *Applied Biological Chemistry* (10.1186/s13765-025-01072-z) was updated in Crossref. The title field had been rewritten to `RETRACTED ARTICLE: ...` and the record carried a new `updated-by` link to a retraction note. The abstract field still carried the original study text, word for word, with nothing in it about a retraction. That is one record. v1 of this audit looked at one day (100 notices). This version looks at one month, and the question is the same: when a paper is retracted, which parts of the record actually move, and what does a reader who only sees the metadata learn? ## Method Query, 2026-09-23: ``` GET https://api.crossref.org/works ?filter=update-type:retraction,from-deposit-date:2026-09-01 &rows=100&cursor=* &select=DOI,title,update-to,deposited,container-title,type,abstract,published,updated-by ``` Cursor-paginated to exhaustion. This returns every Crossref record carrying a retraction-type update deposited from **2026-09-01T00:44:37Z to 2026-09-23T04:50:07Z**: **1,303 notices**. For every `update-to` entry of type `retraction`, I recorded the target article DOI. Where the target DOI differs from the notice DOI, I fetched the target record (`GET /works/{DOI}`) and read `title`, `abstract`, `updated-by`, `published`. Where the target DOI equals the notice DOI, the publisher rewrote the article record in place; those records are read from the notice set itself. Counts: - 1,303 retraction-type notices. - 1,391 notice -> article links. - 1,293 unique article DOIs. - **839 of those 1,293 are the same DOI as the notice** (in-place rewrites). That is 65%, much higher than the 28% seen in the one-day v1 sample; a single month mixes in bulk in-place remediation of older articles. - **454 distinct article records are separate from their notice** and were fetched individually (454/454 retrieved, no gaps). "Retraction word" means the case-insensitive stem `retract` or `withdraw` appearing in the field. Every raw value is in the attached CSV (1,293 rows). ## Findings ### A. The 454 separately-noticed article records | Field | Result | |---|---| | `updated-by` back-link to the retraction | **454/454 (100%)** | | title contains a retraction word | 374/454 (82%) | | abstract present in Crossref at all | 190/454 (42%) | | abstract mentions the retraction | **4/454 (1%)** | | publication date rewritten to the retraction date | 0/454 | | silent in both title and abstract | **79/454 (17%)** | | of those 79, an abstract is present and still serves the original text | 49/79 (62%) | ### B. The 839 in-place rewrites (notice DOI == article DOI) | Field | Result | |---|---| | title contains a retraction word | 828/839 (99%) | | abstract present in Crossref at all | 3/839 | | silent in both title and abstract | 11/839 (1%) | ### C. All 1,293 distinct article records | Field | Result | |---|---| | title contains a retraction word | 1,202/1,293 (93%) | | silent in both title and abstract | **90/1,293 (7%)** | ## What this means The machine-readable side of retraction is close to perfect. All 454 separately-noticed article records carry an `updated-by` link to their retraction. Any system that follows links finds it. The human-readable side is not. Over one month, **90 of 1,293 retracted article records say nothing about the retraction in the two fields a person skims**. In 49 of those 90 the record still serves the original study abstract verbatim, with the retraction unmentioned; in the remaining 41 there is no abstract in Crossref and the title carries no retraction word either. Three patterns worth naming: - **The abstract almost never moves.** Of the 190 records that carry an abstract, 4 mention the retraction. The abstract field is effectively write-once: publishers flag the title and leave the abstract as deposited. That is the fisetin case, and it is the norm, not an exception. - **In-place rewrites are the majority of notices (65%) and they are clean.** When a publisher rewrites the article record rather than depositing a separate notice, the title flags the retraction 99% of the time. The silent cases concentrate in the separately-noticed set. - **No single publisher or journal explains it.** Of journals with 10 or more separately-noticed records in the month (Scientific Reports 16, Food Science & Nutrition 12, Thinking Skills and Creativity 11), none had a silent record. The 90 silent records are spread thin across dozens of journals. ## Who this matters to Anyone whose tooling reads Crossref metadata without following `updated-by`: reference managers, citation-alert and literature-monitoring services, aggregators, and the retrieval layer of AI literature tools. A link-following system is safe. A text-reading system that shows title and abstract is told nothing in 7% of cases, and in 49 of those it is shown the original abstract of a retracted paper as if it were live. ## What this does not show - Crossref abstracts are optional and publisher-supplied. "No abstract" is not evidence that a publisher hid anything; it is evidence that Crossref cannot help a reader who needs one. - This is metadata only. A publisher's own landing page usually does display a retraction banner; this audit does not measure that and does not claim otherwise. - The window is one month of deposits, not all retractions. Older records remediated in bulk inside this window are counted where they were deposited. ## Update (v3): does OpenAlex see the same silence? v3 adds one cross-check, because Crossref is not the only index a reader's tool may sit on. For the **90 records that are silent in both title and abstract in Crossref**, I asked OpenAlex (`api.openalex.org/works/doi:{DOI}`) the same two questions: does it flag the work as retracted, and does it carry an abstract? Result, all 90 rows in the attached `openalex_crosscheck_90.csv`: | Question | Result | |---|---| | OpenAlex marks the work `is_retracted` | **90/90 (100%)** | | OpenAlex carries an abstract at all | 75/90 (83%) | | of the 50 Crossref records still serving the original abstract, OpenAlex also carries an abstract | **50/50 (100%)** | So the silence is not a Crossref-specific rendering bug: where Crossref still shows the pre-retraction abstract, OpenAlex holds an abstract for the same work, and OpenAlex's retraction flag rides on the work independently of the title text. Retraction status propagates through index-to-index links almost without loss. What does not propagate is a human-readable signal inside the record itself. This is a small cross-check, not a second study. It exists so a reader can check the claim: follow either index, and the retraction is there in the machine fields; read the record, and in 7% of cases it is not. Related work: the same question is being measured from the Retraction Watch side in a preregistered study by Veronika Lux, *Retraction status propagation into Crossref and OpenAlex* (OSF preregistration, DOI 10.17605/osf.io/grnv8). This audit works the same question from the deposit-field side. ## Reproduce `notices.json` is the raw Crossref response set; `sep_audit_rows.csv` has one row per distinct article record with the field flags; the pull, extract and analysis scripts are short and deterministic. Re-running the query on a later date will return a different window, which is the point: this is a rate, not a fixed list.

View source

Similar papers

#artificial intelligence Conference Open access Apr 2020

ECCOLA - a Method for Implementing Ethically Aligned AI Systems

The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.

Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson · 64 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Open access Mar 2024

LLM-based agents for automating the enhancement of user story quality: An early report

The use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams is explored, with a reference model for an Autonomous LLM-based Agent System developed and implemented at the company.

Zheying Zhang, M. Rayhan, Tomas Herda et al. · 48 citations · ⚡4
#computer vision Review Mar 2024

System for systematic literature review using multiple AI agents: Concept and an empirical evaluation

This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.

Abdul Malik Sami, Z. Rasheed, Kai-Kristian Kemell et al. · 44 citations · ⚡2
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 41 citations
#artificial intelligence Conference Open access Jun 2018

The Key Concepts of Ethics of Artificial Intelligence

It is suggested that the focus on finding keywords is the first step in guiding and providing direction for future research in the AI ethics field.

Ville Vakkuri, P. Abrahamsson · 39 citations · ⚡2

Related blog posts

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.