# Is AI poisoning the scientific literature? Our comment in Nature

*2025-07-08 — note*


For the past few years, [Sadiq Jaffer](https://toao.com) and I been working with our colleagues in
[Conservation Evidence](https://anil.recoil.org/projects/ce) to do [analysis at scale](https://anil.recoil.org/papers/2024-ce-llm) on the
academic literature. Getting local access to millions of fulltext papers has not
been without drama, but made possible thanks to huge amounts of help from our
[University Library](https://www.lib.cam.ac.uk/) who helped us navigate our
relationships with scientific publishers. We have just **[published a comment
in Nature](https://rdcu.be/evkfj)** about the next phase
of our research, where are looking into the impact of AI advances on evidence synthesis.

<a href="https://rdcu.be/evkfj"> <figure class="image-center"><img src="/images/davidparkins-ai-poison.webp" alt="AI poisoning the literature in a legendary cartoon. Credit: David Parkins, Nature" title="AI poisoning the literature in a legendary cartoon. Credit: David Parkins, Nature" loading="lazy" srcset="/images/davidparkins-ai-poison.768.webp 768w, /images/davidparkins-ai-poison.640.webp 640w, /images/davidparkins-ai-poison.480.webp 480w, /images/davidparkins-ai-poison.320.webp 320w, /images/davidparkins-ai-poison.1024.webp 1024w"><figcaption>AI poisoning the literature in a legendary cartoon. Credit: David Parkins, Nature</figcaption></figure> </a>


Our work on literature reviews led us into assessing methods for [evidence
synthesis](https://royalsociety.org/news-resources/projects/evidence-synthesis/)
(which is crucial to rational policymaking!) and specifically about how recent advances in AI may
impact it.  The current methods for [rigorous systematic literature review](https://en.wikipedia.org/wiki/Systematic_review) are expensive and slow, and authors are already struggling to keep up with the [rapidly expanding](https://ourworldindata.org/grapher/scientific-and-technical-journal-articles?time=latest)
number of legitimate papers. Adding to this, [paper retractions](https://retractionwatch.com/2025/) are increasing near
[exponentially](https://doi.org/10.1038/d41586-023-03974-8) and already
systematic reviews [unknowingly cite](https://retractionwatch.com/the-retraction-watch-leaderboard/top-10-most-highly-cited-retracted-papers/)
retracted papers, with most remaining uncorrected even a year (after notification!)

This is all made much more complex as LLMs are flooding the landscape with
convincing, fake manuscripts and doctored data, potentially overwhelming our
current ability to distinguish fact from fiction.  Just this March, the [AI
Scientist](https://sakana.ai/ai-scientist/) formulated hypotheses, designed and
ran experiments, analysed the results, generated the figures and produced a
manuscript that [passed human peer
review](https://sakana.ai/ai-scientist-first-publication/) for an ICLR
workshop! Distinguishing genuine papers from those produced by LLMs isn't just
a problem for review authors; it's a threat to the very foundation of
scientific knowledge. And meanwhile, Google is taking a different tack with a
collaborative [AI co-scientist](https://research.google/blog/accelerating-scientific-breakthroughs-with-an-ai-co-scientist/) who acts as a multi-agent assistant.
 
So the landscape is moving _really_ quickly! Our proposal for the future of
literature reviews builds on our desire to move towards a more regional,
federated network approach. Instead of having giant repositories of knowledge
that [may be erased unilaterally](https://en.wikipedia.org/wiki/2025_United_States_government_online_resource_removals),
we're aiming for a more bilateral network of "living evidence databases".
Every government, especially those in the Global South, should have the ability to build their
own "[national data libraries](https://anil.recoil.org/notes/uk-national-data-lib)" which represent the body
of digital data that affects their own regional needs.

This system of living evidence databases can be incremental and dynamically
updated, and AI assistance can be used as long as humans remain in-the-loop.
Such a system can continuously gather, screen, and index literature,
automatically remove compromised studies and recalculating results.  We're
working on this on multiple fronts this year; ranging from the computer science
to figure out the distributed-nitty-gritty [^1], over to working with the
[GEOBON folk](https://anil.recoil.org/notes/nas-rs-biodiversity) on global biodiversity [data
management](https://www.tunbury.org/2025/07/02/bon-in-a-box/), and continuing
to drive the core LED design at Conservation Evidence. It feels like a

Read our [Nature Comment piece](https://www.nature.com/articles/d41586-025-02069-w) ([comment on LI](https://www.linkedin.com/posts/anilmadhavapeddy_will-ai-speed-up-literature-reviews-or-derail-activity-7348317711002705920-Y5UT?rcm=ACoAAAB0Kb0BNo1v6ylsGU2NtPa95mj-w1VcaJA)) to learn more about how we think we can safeguard evidence synthesis against the rising tide of "AI-poisoned literature" and ensure the continued integrity of scientific discovery. As a random bit of trivia, the incredibly cool artwork in the piece was drawn by the legendary [David Parkins](https://www.davidparkins.com/), who also drew [Beano](https://www.beano.com/) and [Dennis the Menace](https://en.wikipedia.org/wiki/Dennis_the_Menace_and_Gnasher)\!


[^1]: My instinct is that we'll end up with something [ATProto based](https://arxiv.org/abs/2402.03239) as it's so convenient for [distributed system authentication](https://www.tunbury.org/2025/04/25/bluesky-ssh-authentication/).
Synopsis: Nature comment on AI-generated paper threats to evidence synthesis proposing federated living evidence databases with human-in-loop review.
Words: 546
DOI: 10.59350/pbxew-d2j78

## Related

- [2025 Advent of Agentic Humps: Building a useful O(x)Caml library every day](https://anil.recoil.org/notes/aoah-2025) (note, 2025-12-26)
- [Dear ACM, you're doing AI wrong but you can still get it right](https://anil.recoil.org/notes/acm-ai-recs) (note, 2025-12-22)
- [On the path to the UK/India AI Summit with OpenUK and the ATI](https://anil.recoil.org/notes/path-to-uk-india-ai-summit) (note, 2025-11-11)
- [Royal Society's Future of Scientific Publishing meeting](https://anil.recoil.org/notes/rs-future-of-publishing) (note, 2025-07-14)
- [Will AI speed up literature reviews or derail them entirely?](https://anil.recoil.org/papers/2025-ai-poison) (paper, 2025-07-01)
- [What I learnt at the National Academy of Sciences US-UK Forum on Biodiversity](https://anil.recoil.org/notes/nas-rs-biodiversity) (note, 2025-06-06)
- [Careful design of Large Language Model pipelines enables expert-level retrieval of evidence-based information from syntheses and databases](https://anil.recoil.org/papers/2024-ce-llm) (paper, 2025-05-01)
- [Thoughts on the National Data Library and private research data](https://anil.recoil.org/notes/uk-national-data-lib) (note, 2025-02-17)
- [Conservation Evidence Copilots](https://anil.recoil.org/projects/ce) (project, 2024-01-01)

---
Canonical: https://anil.recoil.org/notes/ai-poisoning
Type: note
License: CC BY 4.0 <https://creativecommons.org/licenses/by/4.0/>
Tags: evidence, llms, ai, federation, networks
