How AI Research Tools Slip Retracted Papers Into Your Thesis
Generative AI cannot tell a retracted paper from a sound one, or a predatory journal from a real one. Both end up in bibliographies, and examiners check the record.
The paper was real, and that was the problem
A fabricated citation is a paper that never existed. A retracted citation is worse in one respect: the paper did exist, was published, was cited, and then was withdrawn because the data were wrong, the analysis was flawed, or the work was fraudulent. A generative AI tool retrieving the literature has no reliable way to know the difference. It returns the retracted paper alongside sound ones, with no warning attached.
The scale is documented. One team submitted 217 retracted papers to an AI tool and asked it to evaluate quality. The tool gave 190 of them high quality scores. Retraction notices posted by journals do not stop continued citation, and the AI systems trained on the published record inherit the citations without inheriting the retractions. A 2025 review found that more than half of a set of retracted articles maintained field citation ratios above one, meaning they kept influencing the literature after withdrawal.
For a thesis candidate this is a verification problem hiding inside a real citation. The reference passes any check that only confirms existence. The DOI resolves. The journal is indexed. The paper is genuine and discredited, and an examiner who knows the field will know it was pulled.
Why AI cannot detect a retraction
The mechanics are straightforward. A generative model is trained on text that existed at a point in time, and a paper's citations accumulate over years while its retraction happens at a single later moment. The training data contain thousands of references to the paper from before it was withdrawn and few or none of the retraction notice. The model learns the paper as a normal, well-cited source.
A pragmatic trial tested nine freely available generative AI tools on whether they would answer questions without citing retracted literature, drawing the retracted articles from the Retraction Watch database. The tools cited retracted work without warning. This is not a tuning failure that a better prompt fixes. It is structural: the model has no live connection to the retraction registry, so it cannot know that a heavily cited paper was discredited after the training cutoff.
The result is that the most-cited retracted papers are the most likely to surface. A paper that was influential before withdrawal generated many citations, which means the model saw it often and ranks it as authoritative. The very prominence that should make a retraction newsworthy is what makes the AI most confident in recommending it.
A candidate in biomedical research is especially exposed, because the fields with the most retractions, including computational biology, neuroscience, and healthcare sciences, are also the fields where deep research tools are most heavily used. The overlap is not coincidental. High-volume publishing produces both the heavy AI use and the retractions.
The predatory journal problem
Retracted papers are one failure. Predatory journals are another, and AI cannot detect them either. A predatory journal exists to collect fees from authors while providing little or no peer review. Its articles have made their way into Google Scholar, Scopus, and occasionally Medline, which means a tool retrieving the literature can return them as if they were peer-reviewed sources.
The tool has no model of journal quality. It sees a formatted citation to a published article and treats it as equivalent to a citation from a rigorous venue. A candidate accepting the output inherits sources that an examiner will recognise as predatory on sight, because examiners know the journals in their field and which ones are not real. A reference to an obscure journal that the examiner has never encountered is a flag, and the candidate who cannot defend the source loses ground.
The check is one a human can do and a generative tool cannot: confirm the journal exists, is indexed in a legitimate database, and is not on a predatory list. AI tools also enable the production of low-quality papers that feed the predatory ecosystem, which means the volume of questionable sources is growing faster than any individual candidate can track. Verification against a quality registry is the only scalable filter.
What this means in the examination room
An examiner who finds a retracted paper or a predatory journal in your bibliography draws a conclusion about your process. The conclusion is not that you are dishonest. It is that you did not verify your sources against the academic record, which is a different and still damaging finding. The bibliography is evidence of how carefully you read, and a discredited source in it undercuts that evidence.
A cited retracted paper is one of the most serious integrity issues an editor catches at desk review, and it is one of the easiest to miss, because the paper was valid when published. The same is true for a thesis. The candidate who built a chapter on a finding that was later retracted has an argument resting on withdrawn evidence, and an examiner who knows the retraction will press on exactly that point.
The defensible position is to have checked. A candidate who can say "I verified every source against Retraction Watch and the indexing databases" has done the work the examiner is testing for. The candidate who relied on a deep research tool's output without verification has, in effect, outsourced a judgment the tool cannot make.
How Thesisroom catches this
Thesisroom's citation guard verifies every bibliography entry against Crossref, OpenAlex, Semantic Scholar, Retraction Watch, and DOAJ. Retraction Watch is the registry that catches a withdrawn paper, and DOAJ is the directory that distinguishes legitimate open-access journals from predatory ones. An entry that resolves to a retracted paper is flagged, not silently passed, because the verification includes the registry the generative tool cannot reach.
The output is a traffic-light list: verified, flagged, or unverifiable. A flag on a retracted paper tells you the source was withdrawn so you can remove the claim or find a sound replacement. When a flagged source needs replacing, Thesisroom's source finder locates a verified alternative from the same registries. The tool does not rewrite your references to hide a retraction or fabricate a clean substitute. Every flag points to a checkable fact, in line with the integrity policy.
Frequently asked questions
Can AI tools tell if a paper has been retracted?
No. Generative models are trained on text from before a retraction occurs, so they learn a withdrawn paper as a normal, well-cited source. They have no live connection to retraction registries. A trial of nine AI tools found they cited retracted literature without warning, drawing on papers listed in the Retraction Watch database.
Why are the most influential retracted papers the most dangerous?
A paper that was influential before withdrawal accumulated many citations, so the model saw it often and ranks it as authoritative. The prominence that should make a retraction newsworthy is exactly what makes the AI most confident in recommending it. More than half of one set of retracted articles kept high citation rates after withdrawal.
Can AI detect a predatory journal?
No. AI tools have no model of journal quality and treat a citation from a predatory journal as equivalent to one from a rigorous venue. Predatory articles reach indexes including Google Scholar and Scopus, so a retrieval tool can return them as if peer-reviewed. Verifying the journal against a quality database is a human check the tool cannot make.
What happens if my thesis cites a retracted paper?
An examiner who knows the retraction will conclude you did not verify your sources, and any argument resting on the withdrawn finding becomes vulnerable. A retracted citation is among the most serious integrity issues caught at review and one of the easiest to miss, because the paper was valid when first published.
Which fields face the highest retraction risk?
Computational biology, neuroscience, and healthcare sciences show the most retractions, and they are also where deep research tools are most heavily used. The overlap is structural: high-volume publishing drives both the AI adoption and the retractions, so candidates in these fields face elevated exposure to discredited sources.
How do I verify my sources are not retracted or predatory?
Check each citation against a retraction registry such as Retraction Watch and confirm the journal is indexed in a legitimate database and absent from predatory lists. Thesisroom's citation guard runs both checks across your full bibliography, flagging retracted papers and unverifiable journals so you can act before submission.
One thing to do before you submit
Run your full bibliography against a retraction registry and a journal index, not just a DOI resolver. Existence is not quality, and a deep research tool cannot tell the difference between a sound paper and a withdrawn one. Thesisroom's citation guard checks every entry against Retraction Watch and DOAJ alongside the standard registries, so a retracted paper or a predatory journal surfaces as a flag in your hands rather than as a question in the viva.