Open research, August 2026
Science has no Shepard's: what happens when the registers of retraction disagree
Legal research has known since 1873 whether an authority still stands. Scientific research does not. We built an open ontology for the integrity status of the scholarly record and measured it against four public registers of retraction. They disagree with each other about 80 per cent of the time, one of them systematically records the announcement of a retraction as retracted research, and the field's central metadata vocabulary contains two different misspellings of the word retraction. Every number below is computed from public data and regenerable.
The short version
- The artefact: an open OWL ontology, a SKOS status vocabulary and a three-layer SHACL suite for the integrity status of scholarly works, plus a 3.19 million triple instance graph across four registers. CC BY 4.0 and MIT.
- One organisation, two datasets, 27.6 per cent disagreement: Crossref acquired Retraction Watch in September 2023 and publishes both. They agree on 72.42 per cent of the retracted DOIs between them.
- The category error: OpenAlex flags 94.5 per cent of retraction notices as retracted research, confirmed independently at 95.95 per cent against Europe PMC. Crossref does this for 0.91 per cent of the same notices, Europe PMC for 0.32 per cent. It is a fixable design defect, not an inherent difficulty.
- Four registers, 19.24 per cent agreement: of 137,243 DOIs asserted retracted by at least one register, only 26,407 are asserted by all four, and 43.09 per cent rest on a single register's say-so.
- Nothing propagates: 43,683 citations to the 400 most-cited retracted works post-date their retraction. The Wakefield paper alone has 1,171, none carrying a machine-readable warning.
- A null result we report rather than bury: OpenAlex almost never holds a retracted paper unflagged, 22 cases in 20,779 probed. The failure is over-flagging, not under-flagging.
The problem legal publishing solved in 1873
In 1873 a salesman for a Chicago legal publisher named Frank Shepard began printing gummed labels. Each listed the later cases that had cited an earlier one, annotated with single-letter codes: this decision overruled, that one criticised, another modified, another simply applied. Lawyers pasted them into the margins of their case reports. The product was called Shepard's Adhesive Annotations, and it solved a problem that had been quietly poisoning legal practice. You could read a case, find it persuasive, cite it, and never learn that a higher court had gutted it three years earlier.
That idea became the citator. A lawyer today runs a citation through Shepard's and gets a status signal before reading a word of the opinion. Citing overturned authority is not merely embarrassing, it is malpractice, and the tooling exists precisely so that it cannot happen by accident.
Science has no citator. It has registers instead, and the registers do not agree with each other. We measured how much, using only open sources, and published the ontology, pipeline and data. What follows is not an argument that the scholarly record is in poor shape. It is a set of counts.
One organisation, two datasets, 27.6 per cent disagreement
In September 2023 Crossref acquired the Retraction Watch database from the Center for Scientific Integrity, paying an initial fee of 175,000 US dollars plus 120,000 a year, and made it openly available. Crossref also operates its own mechanism for recording corrections: publishers register an update-to assertion against a DOI, stating that this record retracts, corrects or raises concern about that one.
So Crossref now publishes two authoritative statements about which papers have been retracted. They are the same organisation's data. Here is how far they overlap.
| Asserted retracted by Retraction Watch | 60,247 |
| Asserted retracted by Crossref update-to | 78,907 |
| Asserted by both | 58,449 |
| Asserted by at least one | 80,705 |
| Agreement | 72.42% |
Just under 28 per cent of the retracted record is asserted by only one of the two. In one direction, 1,798 papers that Retraction Watch records as retracted carry no retraction assertion in Crossref metadata. In the other, 20,458 papers that Crossref metadata marks as retracted are absent from Retraction Watch entirely.
Neither register is wrong, exactly. They were built for different purposes, one by curators reading notices, the other by publishers depositing metadata. But nothing reconciles them, and nothing in the data model of either can express the sentence "the other register disagrees with me".
The corrective apparatus recorded as corrupted literature
A retraction notice is the document that announces a paper has been withdrawn. It is a valid, standing, citable part of the scholarly record. It is the immune response, not the infection.
OpenAlex, the open scholarly graph that increasingly substitutes for Scopus and Web of Science in research tooling, flags 94.5 per cent of Retraction Watch notice DOIs as retracted. Of the DOIs OpenAlex uniquely considers retracted, 74.5 per cent are notice DOIs. In total 40,430 documents whose only role is to announce a retraction, and which were never themselves retracted papers, are recorded as retracted research.
The clearest example is one of the most scrutinised papers of the decade. In 2023 Nature retracted the claim of near-ambient superconductivity in a nitrogen-doped lutetium hydride. The retraction note is DOI 10.1038/s41586-023-06774-2. OpenAlex records that retraction note as retracted.
It would be unfair to rest that on Retraction Watch's own notice column, so we added a fourth register that identifies notices independently. Europe PMC carries two distinct MEDLINE publication types: "Retracted Publication" for the withdrawn paper, and "Retraction of Publication" for the notice announcing it. Taking Europe PMC's 19,803 notice DOIs and asking each register what it says about them:
| Register | Notices it treats as retracted research |
|---|---|
| OpenAlex | 19,001 (95.95%) |
| Crossref | 181 (0.91%) |
| Retraction Watch | 70 (0.35%) |
| Europe PMC itself | 64 (0.32%) |
Two independent measurements, one against Retraction Watch's notice column at 94.5 per cent and one against Europe PMC's publication types at 95.95 per cent, agree. This is not a partial defect at the margins. OpenAlex flags essentially every retraction notice it holds as retracted research, which is the signature of a join performed on the wrong key.
Three registers keep the categories apart and one does not. That settles the question of whether this is hard. It is not hard. It is a design defect, and a fixable one. The consequence is not academic: anything filtering on the retracted flag, whether an integrity dashboard, a bibliometric study or a language model retrieving evidence, receives the corrective apparatus of science mixed into the corrupted literature with no way to separate them.
We should also report a hypothesis that failed. We expected the opposite problem, that OpenAlex would frequently hold a retracted paper without flagging it. It does not. Of 20,779 Retraction Watch DOIs probed and found in OpenAlex, only 22 were unflagged, a rate of 0.11 per cent. Where OpenAlex has the paper it almost always knows. The failure runs the other way.
Adding Europe PMC makes the overall disagreement worse rather than better. Across all four registers, 137,243 DOIs are asserted retracted by at least one. Only 19.24 per cent are asserted by all four, and 43.09 per cent rest on a single register's say-so.
The vocabulary itself is uncontrolled
Crossref's update-type field is how the scholarly record states what kind of correction has occurred. The schema (common5.4.0.xsd) declares it a closed enumeration of 12 values on a required attribute. The live index held 34 distinct values on 16 August 2026, of which 22 fall outside that enumeration, across 365 records. Among them: retration, a misspelling of retraction; retracion, a different misspelling; Retraction and Erratum, case variants that are invalid because the type is case-sensitive; expression-of-concern hyphenated alongside expression_of_concern; 68818, a bare integer; and this_is_some_update_23, which is test data sitting in the production index.
The argument here is not about prevalence. 365 records out of 475,593 is 0.077 per cent, and anyone who objects "so what" to that number is right to. The finding is about governance: the schema declares a closed list on a required attribute, and a string reading this_is_some_update_23 is in the production index regardless. Validation is not enforced at deposit time on this path.
An earlier version of this page reported 19 distinct values. That figure was wrong and is withdrawn. Our pipeline harvests by querying 13 named update-types, so it could only ever observe values co-occurring on records those queries returned; strings like this_is_some_update_23 were unreachable. Crossref's facets enumerate values more completely than a seeded harvest but undercount volumes badly, and the two need to be used for opposite purposes. The correction is recorded in the repository's build report rather than quietly applied.
Each is checkable in about ten seconds. The misspelling retration belongs to DOI 10.1016/j.cie.2010.04.003, published by Elsevier, and the title of that paper literally begins with the word RETRACTED. A human reading the title knows. A machine filtering on update-type retraction does not, because the status string is misspelled. The bare integer belongs to DOI 10.3892/etm.2024.12720, published by Spandidos.
This is the same class of defect we have found in every register examined in this series: a letter O typed in place of a zero inside a Legal Entity Identifier in the European insurance register, checksum-invalid identifiers filed with the SEC in the United States fund universe. Registers are written by people, validation is optional, and what is not validated is eventually wrong.
A related structural defect: in 25.73 per cent of Crossref's retractive assertions the notice is registered against the same DOI as the paper it retracts. The notice therefore has no independent identity. It cannot be cited, linked to, counted or pointed at. The record contains the fact of a retraction but not a retraction document. If you wanted to build a citator for science, this is a layer you would need and largely do not have.
What a citator would actually flag
Everything above concerns whether the record knows a paper is retracted. The citator question is different: does anything warn the people who cite it?
Using citation edges from OpenCitations we populated the propagation layer for the 400 most-cited retracted works and graded every citation against the retraction date. A citation that post-dates the retraction is graded severe. One that pre-dates it is graded caution, because the citing author could not have known, but the reader still cannot tell that the foundation has been withdrawn.
Of 176,623 citations across 391 works, 43,683 post-date the retraction. That is 29.28 per cent of all dated citations, spread across 42,784 distinct citing works, every one of which would carry a warning in any system built the way Shepard's is built. The works generating the most severe signals are not marginal.
| Retracted work | Retracted | Citations after |
|---|---|---|
| PREDIMED Mediterranean diet trial, NEJM | 2018 | 1,251 |
| Wakefield, Lancet, retracted for falsification of data | 2010 | 1,171 |
| Visfatin, Science | 2007 | 1,120 |
| Surgisphere COVID-19 paper, NEJM | 2020 | 833 |
The second row is the paper that launched the modern anti-vaccination movement, retracted in February 2010 for falsification of data. It has been cited 1,171 times since, and not one of those citations carries a machine-readable warning, because no layer of the open scholarly record emits one. This is a sample of the most-cited works rather than a census, and the sampling is biased in ways documented in the repository: the largest citation lists were also the ones most likely to fail to download, so the true figure is higher.
Why the existing ontologies will not do
The natural objection is that this is a data quality problem rather than a modelling problem. We disagree, and the reason is visible in the vocabularies themselves.
The SPAR ontologies are the serious prior art, built largely by Silvio Peroni and colleagues at Bologna. We fetched them and searched them rather than assuming. They do model retraction: CiTO defines cito:retracts and cito:isRetractedBy, FaBiO defines fabio:Retraction, fabio:RetractionNotice and fabio:hasRetractionDate, and PSO defines pso:retracted-from-publication. A practical trap for implementers: CiTO declares its prefix with a hash but defines its terms with a slash, so IRIs built from the prefix declaration do not exist.
What none of them has, verified by searching CiTO, FaBiO, PSO, PRO, SCoRO and DEO, is any of the following: expression of concern, which does not occur as a string in any of the six files; reinstatement, so retraction is modelled as terminal even though registers record works being cleared; partial retraction; removal as distinct from retraction; any way to say that two registers disagree; any way to say that a register holds the record and stays silent; and any notion of a signal propagating to citing works.
There is also a hard logical defect. In FaBiO version 2.3, fabio:hasRetractionDate is declared an owl:FunctionalProperty, while fabio:hasCorrectionDate is not. A functional property may take only one value per subject. So two registers reporting different retraction dates for the same paper does not merely go unmodelled: it makes the graph logically inconsistent. In the incumbent vocabulary, register disagreement about a retraction date is a contradiction rather than an observation, and the defect is specific to retraction.
The deepest of these is the disagreement gap, and it is structural rather than an oversight. These vocabularies model retraction status as a property of a work. Once you write it that way a graph can hold only one answer, and the fact that sources conflict becomes unrepresentable. You are forced to pick a winner before you have modelled the evidence.
What the ontology models
The Scholarly Record Ontology starts from a different commitment. Retraction status is not a property of a work. It is a dated claim, made by a named register, retaining that register's own words.
Every claim is reified and attributed. An IntegrityAssertion names exactly one Register, one work, one status and, where available, one date. Two assertions about the same work may carry different statuses without contradiction at the data layer, because each is attributed to its own source. An unattributed integrity claim is not usable evidence, and the SHACL layer enforces that.
The register's own words are kept verbatim. The raw status string is stored alongside the normalised one. Normalising retration into retracted and discarding the original destroys the evidence that the vocabulary is broken. This is the difference between a graph that cleans data and a graph that can be used to audit the sources it was built from.
Disagreement is first-class. RegisterDisagreement carries a kind, an asserting register and, critically, a silent register. Silence is a position: a graph that does not flag a retracted work will be queried as though the work were sound. Nothing in the existing scholarly vocabularies can express that.
Notices are distinguished from the works they correct. CorrectiveNotice is a distinct class, and where a publisher registers a correction against the original DOI the assertion is marked self-referential rather than silently collapsed. This is precisely the distinction whose absence produces the OpenAlex defect measured above.
Signals propagate, and carry provenance. A PropagationSignal grades a citing work by the status of what it cites and by whether the citation post-dates the assertion, and must name the assertion it derives from. A signal that cannot be traced to a named assertion must not be shown to a reader.
The same discipline was proved first on the United States fund universe in our Investment Fund Ontology and then on European insurers in the Insurance Register Ontology. The scholarly record is its third instantiation, which is itself the finding: reified, source-attributed identifier and status assertions transfer across entirely unrelated register fabrics without redesign.
A methodological lesson worth stating
Partway through harvesting, OpenAlex began returning HTTP 429 with the message "Insufficient budget. This request costs $0.0001 but you only have $0 remaining." OpenAlex now meters its API at roughly 1,000 requests, or ten cents, per day free. The harvest stopped at 107,200 of 134,094 records and was completed the following day. That is worth knowing on its own, because the argument that OpenAlex is the open replacement for Scopus is weaker when the open replacement is metered.
The interruption produced the most instructive lesson of the exercise. While the data was partial we noted that cursor order follows OpenAlex work ID, which correlates with when records were created, so the subset was not a random sample and its percentages should be treated as indicative. That turned out to matter enormously. The final 26,913 records were 56.6 per cent Retraction Watch notice DOIs and 33.2 per cent Europe PMC notice DOIs, against only 5.1 per cent actual retracted papers. The tail of the harvest was almost entirely notices.
So the notice-conflation figure moved from a provisional 64.2 per cent to a complete 94.5 per cent, and the Europe PMC control from 51.15 per cent to 95.95 per cent. Publishing the partial figure without the caveat would have understated the defect by a third. A partial harvest is not a small version of the whole: where the sort key correlates with how records were created, the missing slice can be exactly the slice that matters.
The corporate symmetry
There is a fact about the ownership of all this worth stating plainly, because it makes the gap concrete rather than abstract. Reed Elsevier acquired Shepard's in 1996 and took full ownership in 1998. Reed Elsevier is now RELX, which also owns Elsevier, the largest publisher of scientific literature, and Scopus, one of the two dominant citation databases. Its LexisNexis Risk Solutions division sells entity resolution, the discipline of deciding when two records refer to the same real-world thing, to banks and governments.
RELX has already taken the citator idea further than most people outside legal publishing realise. Its 2024 Annual Report describes Lexis+ AI as using "the LexisNexis proprietary Retrieval Augmented Generation platform, integrated with advanced Shepard's Knowledge Graph", so that customers can harness Shepard's case law relationship information for "authoritative, complete, and final AI-generated responses".
RELX is more explicit still in its 2025 Annual Report, where a diagram of how the group adds value with generative AI places "Knowledge Graphs" in the grounding layer, along an axis labelled "Decreasing hallucination, irrelevant content, non-attributable content (lack of citations)". The company's own published position is that knowledge graphs and linking are the mechanism by which attributability is restored. The same report is the only place where Elsevier's product language names the technique directly, describing a multi-model approach adapted for specific domains "through hybrid search, knowledge graphs, ontologies, large language model and human expertise-based evaluations".
Set against that, one absence is striking. The word "retract" does not appear anywhere in the 252 pages of the RELX 2025 Annual Report, nor anywhere in its Form 20-F for the same year. Research integrity is disclosed as a formal risk factor, and the detection tooling is described in some detail, but no retraction count, no paper-mill interception figure and no research-integrity metric is published. For a group whose case for its AI products rests on the trustworthiness of its content, that is a conspicuous gap in the reporting rather than in the work.
Read that with the scholarly record in mind. RELX has concluded that grounding a language model on legal content is not sufficient on its own, and that the model must also be wired into a knowledge graph of citation treatment so that it does not confidently cite authority that has been overturned. That is exactly the architecture the scientific record lacks, built by the same company, and shipped as a product. If citation-treatment grounding is necessary to stop a legal AI relying on overturned authority, it is necessary for the same reason to stop a scientific AI relying on retracted findings. The difference is not technical. It is that one market pays for the assurance and the other has never been asked to.
This is not hypocrisy and we are not accusing anyone of it. In law the citator is the product, because clients pay for the assurance that authority still stands. In science the citation database is the product, and the assurance layer has never been the thing anyone was buying. It is also worth recording that on the propagation measure Elsevier performs well: 98.9 per cent of Elsevier-published retracted papers carry a corresponding Crossref assertion. The worst performers in the league table are elsewhere, including publishers with zero per cent coverage across hundreds of retracted papers. The problem is not that any one publisher is negligent. It is that no layer above them reconciles anything, and no vocabulary in use can describe the reconciliation.
Related work
One thing needs saying plainly before anything else in this section. The fact that retraction registers disagree is not a discovery of this study. It is an established research programme, principally that of Jodi Schneider and colleagues. Salami, McCumber and Schneider (2026) compare eleven sources across a union of 83,317 items and find that 92.64 per cent carry some disagreement, with only 41 items recorded as retracted in all eleven; their earlier four-source study found 3 per cent agreement. The 72.42 per cent figure reported above is a two-source special case of that result. Anyone citing this work should cite theirs.
The same group has also already documented the category error for Crossref. Si, Salami and Schneider report that 661 of 925 sampled DOIs, 71.46 per cent, are retraction notices indexed as retracted papers. What this study adds there is the four-register comparison, the Europe PMC control, and the measurement that the conflation is near-total in OpenAlex at 94.5 and 95.95 per cent while being rare in Crossref by our own count at 0.91 per cent. Separately, Hauschke and Nazarovets documented an earlier OpenAlex episode in which granular Crossref metadata collapsed into a single Boolean; that incident was transient and repaired, whereas the conflation measured here is structural and current.
On the modelling side, RIPE-O, the Research Integrity Provenance and Evidence Ontology of Markovic and Indukuri (ISWC 2026), is the closest existing artefact and does model expressions of concern. It is built for the provenance of integrity assessments of clinical trials, carries a single source identifier for Retraction Watch, and does not model disagreement between registers. The RISRS report of 2022 recommended "a taxonomy of retraction categories and corresponding retraction metadata that can be adopted by all stakeholders", and this work is an attempt at part of that.
So what is actually claimed here, and only this: a formal, source-attributed representation in which register silence can be stated; the four states no scholarly vocabulary can express; the functional-property defect above; Europe PMC as a register, which no published comparison study includes; and the propagation layer. The empirical fact of disagreement belongs to the authors named above.
Retraction also has a wider quantitative literature that this sits alongside. Jonas Oppenlaender's How Ten Publishers Retract Research (arXiv:2602.19197, February 2026) analyses 46,087 retractions in the Retraction Watch database and reports retraction rates per publisher: Hindawi at 320.02 per 10,000 published, IEEE at 17.70, Springer Nature at 9.06 and Elsevier lowest of the ten at 3.97. That work measures who retracts. This one measures whether the registers agree that a retraction happened at all, and whether anything propagates the result, which is a different question about the same corpus.
One finding of Oppenlaender's bears directly on the ontology. Of 98 articles reinstated following retraction, 86 were published in Elsevier journals. Reinstatement is therefore a real state that occurs at measurable scale, and Retraction Watch records 160 such cases. It is also a state that no existing scholarly vocabulary can express: the string does not occur in CiTO, FaBiO, PSO, PRO, SCoRO or DEO. A record that models retraction as terminal will continue to carry a warning against work that has been cleared, which is a reputational harm to named authors rather than a modelling inconvenience.
A note on reading Elsevier's low retraction rate. It is equally consistent with a cleaner corpus and with a more conservative retraction practice, and the pairing with the highest reinstatement share is what makes it interesting rather than settled. Our own propagation measurement is the narrower and more defensible claim: where Elsevier does retract, 98.9 per cent of those papers carry a corresponding Crossref assertion.
What we are not claiming
We cannot compare any of this against Scopus or Web of Science, because both are licensed products with no open API for the purpose. That limitation is itself part of the point: the open record is what most downstream tooling actually consumes, and increasingly what language models are trained and grounded on.
The propagation layer is a deliberate sample of the 400 most-cited retracted works, not a census, and its selection is biased. The licence under which the Retraction Watch data is redistributed is not stated on Crossref's documentation page, so the raw data is not committed to the repository, only the code that fetches it. Every caveat of this kind is written up in the repository's build report rather than left for a reader to discover.
What would fix it
A citator for science is not a research problem. Every component already exists. Retraction Watch curates. Crossref receives publisher assertions. Europe PMC keeps the categories straight. OpenAlex and OpenCitations hold the citation edges. What is missing is a layer that treats these as competing sources rather than as a single truth, records where they conflict, and propagates a graded signal to the works that cite them.
That layer needs a vocabulary that can express disagreement. That is what this ontology is for, and it is open, with the pipeline and the graph.
Work with us on this
If you work on research integrity, scholarly infrastructure, publishing metadata or knowledge graphs and want the underlying data, a walkthrough, or this analysis run against your own corpus, we would like to hear from you. If a number here is wrong, tell us with the DOI and we will recheck it against source and credit the correction.
The repository on GitHub (opens in new tab) contains the ontology, the SHACL shapes, six worked SPARQL queries, the full pipeline and a build report listing every caveat.
