# A fully AI-generated paper just passed peer review; notes from our evidence synthesis workshop

*2025-03-12 — note*


Access to reliable and timely scientific evidence is utterly vital for the practise of responsible policymaking, especially with all the turmoil in the world these days. At the same time, the evidence base on which use to make these decisions is rapidly morphing under our feet; the [first entirely AI-generated paper passed peer review](https://sakana.ai/ai-scientist-first-publication/) at an ICLR workshop today.  We held a workshop on this topic of AI and evidence synthesis at [Pembroke College](https://pem.cam.ac.uk) last week, to understand both the opportunities for the use of AI here, the [strengths and limitations](https://anil.recoil.org/papers/2024-ce-llm) of current tools, areas of progress and also just to chat with policymakers from [DSIT](https://www.gov.uk/government/organisations/department-for-science-innovation-and-technology) and thinktanks about how to approach this rapidly moving area.


*(The following notes are adapted from jottings from [Jessica Montgomery](https://www.cst.cam.ac.uk/people/jkm40),
[Sam Reynolds](https://samreynolds.org), [Annabelle Scott](https://ai.cam.ac.uk/people/annabelle-scott) and myself. They are not at all complete, but hopefully useful!)*

We invited a range of participants to the workshop and held it at Pembroke College (the choice of the centuries-old location felt appropriate).
[Jessica Montgomery](https://www.cst.cam.ac.uk/people/jkm40) and [Neil Lawrence](https://inverseprobability.com/) expertly emceed the day, with [Bill Sutherland](https://www.zoo.cam.ac.uk/directory/bill-sutherland), [Sadiq Jaffer](https://toao.com) and [Sam Reynolds](https://samreynolds.org) also presenting provocations to get the conversation going.

<figure class="image-center"><img src="/images/evidence-synth-2.webp" alt="Lots of excellent discussions over Pembroke sarnies!" title="Lots of excellent discussions over Pembroke sarnies!" loading="lazy" srcset="/images/evidence-synth-2.768.webp 768w, /images/evidence-synth-2.640.webp 640w, /images/evidence-synth-2.480.webp 480w, /images/evidence-synth-2.3840.webp 3840w, /images/evidence-synth-2.320.webp 320w, /images/evidence-synth-2.2560.webp 2560w, /images/evidence-synth-2.1920.webp 1920w, /images/evidence-synth-2.1600.webp 1600w, /images/evidence-synth-2.1440.webp 1440w, /images/evidence-synth-2.1280.webp 1280w, /images/evidence-synth-2.1024.webp 1024w"><figcaption>Lots of excellent discussions over Pembroke sarnies!</figcaption></figure>

## Evidence synthesis at scale

[Jessica Montgomery](https://www.cst.cam.ac.uk/people/jkm40) described the purpose of the workshop as follows:

> Evidence synthesis is a vital tool to connect scientific knowledge to areas
> of demand for actionable insights. It helps build supply chains of ideas,
> that connect research to practice in ways that can deliver meaningful
> improvements in policy development and implementation.  Its value can be seen
> across sectors: aviation safety benefitted from systematic incident analysis;
> medical care has advanced through clinical trials and systematic reviews;
> engineering is enhanced through evidence-based design standards. When done
> well, evidence synthesis can transform how fields operate. However, for every
> field where evidence synthesis is embedded in standard operating practices,
> there are others relying on untested assumptions or outdated guidance.
> <cite>\-- [Jessica Montgomery](https://www.cst.cam.ac.uk/people/jkm40), AI@Cam</cite>

One such field that benefits from evidence is [conservation](https://anil.recoil.org/projects/ce), which is what [Bill Sutherland](https://www.zoo.cam.ac.uk/directory/bill-sutherland) and his [team](https://conservationevidence.com) have been working away on for years.  Bill went on to discuss the fresh challenges that AI brings to this field, because it introduces a new element of scale which could augment relatively slow human efforts.

> Scale poses a fundamental challenge to traditional approaches to evidence
> synthesis.  Comprehensive reviews take substantial resources and time. By the
> time they are complete – or reach a policy audience – the window for action
> may have closed.  The Conservation Evidence project at the University of
> Cambridge offers an example of how researchers can tackle this challenge. The
> Conservation Evidence team has analysed over 1.3M journals from 17 languages
> and built a website enabling access to this evidence base.  To support users
> to interrogate this evidence base, the team has compiled a metadataset that
> allows users to explore this literature based on a question of interest, for
> example looking at what conservation actions have been effective in managing
> a particular invasive species in a specified geographic area.
> <cite>\-- [Jessica Montgomery](https://www.cst.cam.ac.uk/people/jkm40), AI@Cam</cite>

The AI for evidence synthesis landscape is changing very rapidly, with a variety of specialised tools now
being promoted in this space. This ranges from commercial tools such as [Gemini Deep Research](https://gemini.google/overview/deep-research/?hl=en) and [OpenAI's deep searcher](https://openai.com/index/introducing-deep-research/), to
research-focused systems such as [Elicit](https://elicit.com), [DistillerSR](https://www.distillersr.com/products/distillersr-systematic-review-software), and [RobotReviewer](https://www.robotreviewer.net). These tools vary in their approach, capabilities, and target users, raising questions about which will best serve different user needs.  RobotReviewer, for example, notes that:

> \[...\] the machine learning works well, but is not a substitute for human systematic reviewers. We recommend the use of our demo as an assistant to human reviewers, who can validate the machine learning suggestions, and correct them as needed. Machine learning used this way is often described as semi-automation.
> <cite>\-- [About RobotReviewer](https://www.robotreviewer.net/about)</cite>

The problem, of course, is that these guidelines will often be ignored by
reviewers who are under time pressure, and so the well established protocols
for systematic reviewers are under some threat.

<figure class="image-center"><img src="/images/evidence-synth-4.webp" alt="Sadiq Jaffer and Sam Reynolds discuss emerging AI systems" title="Sadiq Jaffer and Sam Reynolds discuss emerging AI systems" loading="lazy" srcset="/images/evidence-synth-4.768.webp 768w, /images/evidence-synth-4.640.webp 640w, /images/evidence-synth-4.480.webp 480w, /images/evidence-synth-4.320.webp 320w, /images/evidence-synth-4.1920.webp 1920w, /images/evidence-synth-4.1600.webp 1600w, /images/evidence-synth-4.1440.webp 1440w, /images/evidence-synth-4.1280.webp 1280w, /images/evidence-synth-4.1024.webp 1024w"><figcaption>Sadiq Jaffer and Sam Reynolds discuss emerging AI systems</figcaption></figure>

## How do we get more systematic AI-driven systematic reviews?

[Sadiq Jaffer](https://toao.com) and [Sam Reynolds](https://samreynolds.org) then talked about some of the computing approaches required to achieve a more reliable evidence review base.
They identified three key principles for responsible AI integration into evidence synthesis:
- Traceability: Users should see which information sources informed the evidence review system and why any specific evidence was included or excluded.
- Transparency: Open-source computation code, the use of open-weights models, [ethically sourced](https://www.ibm.com/impact/ai-ethics) training data, and clear documentation of methods mean users can scrutinise how the system is working.
- Dynamism: The evidence outputs should be continuous updated to refines the evidence base, via adding new evidence and flagging [retracted papers](https://anil.recoil.org/notes/ai-contamination-of-papers).

[Alex Marcoci](https://www.cser.ac.uk/team/alex-marcoci/) pointed out his recent work on [AI replication games](https://osf.io/sz2g8/) which I found fascinating. The idea here is that:

> Researchers will be randomly assigned to one of three teams: Machine, Cyborg
> or Human. Machine and Cyborg teams will have access to (commercially
> available) LLM models to conduct their work; Human teams of course rely only
> on unaugmented human skills. Each team consists of 3 members with similar
> research interests and varying skill levels. Teams will be asked to check for
> coding errors and conduct a robustness reproduction, which is the ability to
> duplicate the results of a prior study using the same data but different
> procedures as were used by the original investigator.
> <cite>\-- [Institute for Replication](https://www.sheffield.ac.uk/machine-intelligence/events/i4rs-ai-replication-games)</cite>

These replication games are happening on the outputs of evidence, but the
*inputs* are also rapidly changing with today's announcement of a [fully generated AI papers passing peer
review](https://sakana.ai/ai-scientist-first-publication/). It's hopefully now clear
that AI is a huge disruptive factor in evidence synthesis.

<figure class="image-center"><img src="/images/evidence-synth-3.webp" alt="" title="" loading="lazy" srcset="/images/evidence-synth-3.768.webp 768w, /images/evidence-synth-3.640.webp 640w, /images/evidence-synth-3.480.webp 480w, /images/evidence-synth-3.3840.webp 3840w, /images/evidence-synth-3.320.webp 320w, /images/evidence-synth-3.2560.webp 2560w, /images/evidence-synth-3.1920.webp 1920w, /images/evidence-synth-3.1600.webp 1600w, /images/evidence-synth-3.1440.webp 1440w, /images/evidence-synth-3.1280.webp 1280w, /images/evidence-synth-3.1024.webp 1024w"><figcaption></figcaption></figure>

## The opportunity ahead of us for public policy

We first discussed how AI could help in enhancing systematic reviews.
AI-enabled analysis can accelerate literature screening and data extraction,
therefore helping make the reviews more timely and comprehensive.  The
opportunity ahead of us is to democratise access to knowledge synthesis by
making it available to those without specialised training or institutional
resources, and therefore getting wider deployment in countries and
organisations without the resources to commission traditional reviews.

However, there are big challenges remaining in [gaining access](https://anil.recoil.org/notes/uk-national-data-lib) to published research papers and datasets.
The publishers have deep concerns over AI-generated evidence synthesis, and more generally about the use of generative AI involving their source material. But individual publishers are [already selling](https://theconversation.com/an-academic-publisher-has-struck-an-ai-data-deal-with-microsoft-without-their-authors-knowledge-235203) their content to the highest bidder as part of the [data hoarding wars](https://anil.recoil.org/notes/ai-ietf-aiprefs) and so the spread of the work into pretrained models is not currently happening equitably or predictably.
[Neil Lawrence](https://inverseprobability.com/) called this "competitive exclusion", and it is limiting communication and knowledge diversity.

The brilliant [Jennifer Schooling](https://www.aru.ac.uk/people/jennifer-schooling) then led a panel discussion about the responsible
use of AI in the public sector.  The panel observed that different countries
are taking different approaches to the applications of AI in policy research.
However, every country has deep regional variances in the *application* of
policy and priorities, which means that global pretrained AI models always need
some localized retuning. The "one-size-fits-all" approach works particularly
badly for policy, where local context is crucial to a good community outcome
that minimises harm.

Policymakers therefore need realistic expectations about what AI can and cannot do in evidence synthesis.
[Neil Lawrence](https://inverseprobability.com/) and [Jennifer Schooling](https://www.aru.ac.uk/people/jennifer-schooling) came up with the notion that "anticipate, test, and learn" methods must guide AI deployment in policy research; this is an extension of the "[test and learn](https://public.digital/pd-insights/blog/2024/12/just-what-is-test-and-learn)" culture being pushed by Pat McFadden as part of the Labour plan to [reform the public sector](https://www.gov.uk/government/speeches/reform-of-the-state-has-to-deliver-for-the-people) this year.  With AI systems, [Alex Marcoci](https://www.cser.ac.uk/team/alex-marcoci/) noted that we need to be working with the end users of the tools to scope what government departments need and want. These conversations needs to happen *before* we build the tools, letting us anticipate problems before we deploy and test them in a real policy environment. [Neil Lawrence](https://inverseprobability.com/) noted that policy doesn't have a simple "sandbox" environment to test AI outcomes in, unlike many other fields where simulation is practical ahead of deployment.

[Lucia Reisch](https://www.jbs.cam.ac.uk/people/lucia-reisch/) noted that users must maintain critical judgement when using these
new AI tools; the machine interfaces must empower users towrads enhancing their
critical thinking and encouraging reflection on what outputs are being created
(and what is being left out!).  Lucia also mentioned that her group helps run
the "[What Works](https://whatworksclimate.solutions/about/)" summit, which
I've never been to but plan on attending next it rolls around.

The energy requirements for training and running these large scale AI models
are significant as well, of course, raising questions about the long-term
maintenance costs of these tools and their environmental footprint.  There was
wide consensus that the UK should develop its own AI models to ensure
resilience and sovereignty, but also to make sure that the regional finetuning
to maximise positive outcomes is under clear local control and not outsourced
geopolitically. By providing a single model that combines [UK national data](https://anil.recoil.org/notes/uk-national-data-lib), we would also not waste energy with lots of
smaller training efforts around the four nations.

<figure class="image-center"><img src="/images/evidence-synth-1.webp" alt="Sadiq Jaffer in front of a very old, very fancy and not AI-designed door" title="Sadiq Jaffer in front of a very old, very fancy and not AI-designed door" loading="lazy" srcset="/images/evidence-synth-1.768.webp 768w, /images/evidence-synth-1.640.webp 640w, /images/evidence-synth-1.480.webp 480w, /images/evidence-synth-1.3840.webp 3840w, /images/evidence-synth-1.320.webp 320w, /images/evidence-synth-1.2560.webp 2560w, /images/evidence-synth-1.1920.webp 1920w, /images/evidence-synth-1.1600.webp 1600w, /images/evidence-synth-1.1440.webp 1440w, /images/evidence-synth-1.1280.webp 1280w, /images/evidence-synth-1.1024.webp 1024w"><figcaption>Sadiq Jaffer in front of a very old, very fancy and not AI-designed door</figcaption></figure>

Thanks [Annabelle Scott](https://ai.cam.ac.uk/people/annabelle-scott) for such a stellar organisation job and to Pembroke for hosting and all for
attending, and please do continue the discussion about this [on LinkedIn](https://www.linkedin.com/feed/update/urn:li:activity:7303431795587309569/)
if you are so inclined.
Synopsis: "AI-generated paper passes peer review, sparking discussion on evidence synthesis and AI's role in policymaking."
Words: 1565
DOI: 10.59350/k540h-6h993

## Related

- [AI, science and the UK–EU relationship at the Royal Society](https://anil.recoil.org/notes/rs-eu-ai-science) (note, 2026-04-21)
- [Foundational AI for Ecosystem Resilience workshop](https://anil.recoil.org/notes/foundational-ecosystem-workshop) (note, 2025-12-03)
- [Four Ps for Building Massive Collective Knowledge Systems](https://anil.recoil.org/notes/principles-for-collective-knowledge) (note, 2025-11-23)
- [Careful design of Large Language Model pipelines enables expert-level retrieval of evidence-based information from syntheses and databases](https://anil.recoil.org/papers/2024-ce-llm) (paper, 2025-05-01)
- [The AIETF arrives, and not a moment too soon](https://anil.recoil.org/notes/ai-ietf-aiprefs) (note, 2025-02-28)
- [Thoughts on the National Data Library and private research data](https://anil.recoil.org/notes/uk-national-data-lib) (note, 2025-02-17)
- [Fake papers abound in the literature](https://anil.recoil.org/notes/ai-contamination-of-papers) (note, 2025-02-04)
- [Conservation Evidence Copilots](https://anil.recoil.org/projects/ce) (project, 2024-01-01)

---
Canonical: https://anil.recoil.org/notes/ai-for-evidence-synthesis-workshop
Type: note
License: CC BY 4.0 <https://creativecommons.org/licenses/by/4.0/>
Tags: ce, conservation, ai, llms, evidence
