This is my last week of academic sabbatical! My previous one happened during the pandemic, and so this is the first proper break I've had for a long time. Looking back on it, the pace of change of technology has been so great this year that I'm very grateful to have had the time to ride the wave and keep up. I don't think I could have done so without the space to experiment that a sabbatical gives, and I feel much better equipped for this new agentic world we're heading into.
I'm always surprised by how many people comment that I must be relieved not to be teaching. The opposite is true: I love teaching, and wouldn't be a professor otherwise! What I really didn't miss is the administrative load that gets higher every year. This time around, I'm taking on being an examiner for the 1B CST Tripos, as well as Pembroke committee memberships for the finance and scrutiny and development and engagement committees, and also resuming the CRI management board and the University environmental committee. I've probably lost track of a few others that are in my inbox somewhere, but I'm sure I'll figure it out as term unfolds!
As a committee counterbalance, I wonder if I'll find time to action the Cambridge Green Blue sporting idea this year. John Palfrey also mentioned to me that Harvard have a similar competitive Green cup, so there's precedent for some grassroots action here!
1 Prepping for the societal agentopalypse
What's foremost on my mind is how to help our students prepare for the rapid societal changes happening with agentic AI, and so assembling a good reading list comes first to educate myself. Kathi Fisler has written an excellent paper on teaching CST in the agentic era that's going to be in the November CACM. It's too late to change this year's course, but I'm already preparing a larger reboot of 1A FoCS with this in mind.
1.1 Universities need to shift back to formative discourse

We discussed how different parts of society take on formative versus summative approaches. That is, how much do we want to just achieve an outcome ('fix that pothole!'), and how much do we want a societal function that keeps humans usefully occupied ('you wouldn't send a robot to the gym for you'). Universities are obviously deeply worried about this divide, since many of our assessment regimes have drifted to summative assessments (like MOOCs) and lost track of the importance of formative discourse in teaching.
Cambridge, with its built-in collegiate inefficiency, does still have an edge here due to our residential structure. However, it's no easy thing to be a student these days! Urs listed several reasons for their difficulties in Germany: students hold multiple jobs due to the rising cost of living, they live further out due to urban gentrification where older universities are, and they have to digest more and more course content. Cambridge does reasonably well on the second point with our Colleges, but the cost of living is now very high and the course content is getting more packed.
1.2 My reading list for October
Urs also mentioned he used to work with John Palfrey at Harvard. John mentioned to me recently an entertaining factlet about who decides where the next AI summit is, so I happily discovered he's also a blogger and have added him to my blogroll!

- BiblioTech: Why Libraries Matter More Than Ever in the Age of Google due to our work on E-TAP. I just picked up a second hand copy in the Market Square book stall here!
- Born Digital, which was originally written two decades ago and then refreshed in 2016, so it's a good time to pick up how things are changing now for young people.
- Zoe Jaques (professor of children's literature here) explained at our monthly reading club how hard it is to write a children's novel without adversity. The characters are usually orphans or have absent parents because otherwise the children have no agency to experience a fun story! It's really difficult to construct interesting protopian scenarios as it's far easier to fearmonger AI extinction instead. Therefore, I'm reading more about how to write protopian screenplays to prep some future histories for E-TAP and get us thinking more constructively about where we're going in our research.
- Rob Doubleday pointed me to Storylistening: a theory and practice for gathering narrative evidence that will complement and strengthen, not distort, other forms of evidence, including that from science. Again, bang on our E-TAP ambitions so it's on the list!
- It's pretty clear that GDP is no longer (ever?) the right measure of progress for a society, but it is much less clear what should replace it. Jeremy Adelman recommended to me Diane Coyle's The Measure of Progress, which is now on my list. Cambridge is a deliberately inefficient place in many ways, and students learn better because of it, but the question remains about how we measure (and reward) inefficiency! John Aston observed in our dinner conversation that societal policies tend to favour things we measure, and so it's just as important to figure out alternatives to GDP as it is to criticise the pervasive use of it!
I think this is a solid list to get me going. If you have any further recommendations I'd be delighted to publish a proper reading list...
2 On more natural matters
2.1 Community sensing and grassroots regeneration
Jon Crowcroft dropped me a note about Bristol Nature Telemetry via Laura James. They build privacy-first acoustic sensors that classify birds and bats on the device, so that no recordings leave it. I learnt a while back from Alec Christie and Sam Reynolds that when acoustic monitors record people it's termed human bycatch. I need to talk to Josh Millar to see if we can incorporate this into the distillation process for uNPU embedded training.
I've also finished my Cambridge mirror of OpenStreetMap for Michael Dales to use for habitat mapping and need to clean up the containers for local use. Alastair Tse also told me about Safecast, a radiation monitoring layer from all over the world (but sadly focussed on areas like Ukraine due to the war and Japan in the post-reactor accident world), all built from crowd-sourced measurements. I'm finding more and more these community-run sensor networks that complement what satellites can observe with Tessera, so I've started thinking about how to fuse these together into ground-level embeddings. This is of particular interest after my chat with the Echo Labs folk about fusing bioacoustic foundation models with satellite ones.
Ultimately, the outcome we're looking for here is how normal people can use these products to help them with grassroots level community rebuilding. It's not just big politicians who should sweep policies from the top down! Laura James wrote a great post on her first FLIPIM, the grassroots regeneration conference, which ties into the rewilding the web theme of late. I look forward to attending a future edition of these.

2.2 Good news from the Cairngorms
In more nature related news, our woodland expansion paper came out in the Journal of Applied Ecology this week. I wrote up another mini explainer about how the airborne LiDAR survey we commissioned back in 2023 over the Cairngorms Connect area mapped 2.65 million young trees and showed that deer management is working.
The feedback has been very positive, and several people got in touch asking about how to get access to more UK LiDAR data. This will all be much easier once we port the pipeline to run over Tessera, as it handles a lot of base layer details that can then be fine-tuned with LiDAR data!
2.3 Habitat mapping progress
Michael Dales and I have also been putting our heads together on how best to support David Coomes, Aneesh Naik and James G. C. Ball in their habitat mapping project, now that global v1.1 coverage exists (more on that below). Michael's weeknotes describe the problem very well; his method works well in the relatively data-rich UK (with lots of GBIF occurrences and OpenStreetMap land types), but gives poorer results in data-deficient regions such as Madagascar.
Since Tessera embeddings are in theory globally comparable, he hopes to borrow habitat information from similar regions elsewhere. Our next step is a global run with help from Sadiq Jaffer, and then we'll try to build a habitat labeling pipeline to combine the best of all the methods Aneesh Naik, James G. C. Ball and Michael have all been experimenting with in the past few months.
An interesting adjacent idea about pollen from Alastair Tse also came up this week. He mentioned that the Japanese cedar ("sugi") causes a national hay fever problem in Japan! BBC Future traces this to a 1950s reforestation that didn't use enough tree diversity, and now about 43% of the population have medium to severe symptoms, vs 26% in the UK (!). Japan now wants to plant no-pollen trees, and a 2023 paper used CRISPR/Cas9 to produce no-pollen sugi. It would be a cool project to try to map Tessera and wind models over to pollen occurrences in Japan!
3 Tessera goes global now on source coop as well
3.1 From dClimate's Icechunk to Source Coop Zarr
dClimate published their writeup on how they re-engineered the Tessera v1.1 inference pipeline into something that runs on cloud pipelines rather than our 'supercomputers' (DAWN, Isambard-AI) that the original Cambridge research codebase assumed. They rebuilt our research codebase around Dask for satellite data ingestion, and Ray for embeddings inference, also moving to cloud-native Zarr format like we've been doing.
We cut the wall time for a year of embeddings over a medium-sized area, such as a US state, from days to under an hour. On the global campaign, that meant finishing almost 65% under our already ambitious budget. dClimate, 2026
Their million-dollar-(ish) run is what produced the whole 2017-2025 Icechunk archive we've been working with. Mark Elvers and I have been converting it this week from their Icechunk store into the Zarr v3 layout GeoTessera reads, and the conversion has now finished after about 13.5 days and roughly 125,000 vCPU-hours, costing about $3k on Fargate Spot. The only minor surprise we ran into was that dClimate skipped some tiles with sparse observations, but our scripts were easily adapted to fit these 'holes'.
We've also got access now to Intel's AMX CPUs on their Endeavour cluster, where we're hoping to do global v2 inference. It's still quite a bit slower than a GPU (a single tile takes 6.5-8 minutes on an AMX CPU, vs about 30s on an MI355X) but many optimisations remain. I'll also release a new geotessera 0.11 to make v1.1 the default and read the global dClimate Icechunk store now that we have 2017-2025 all inferred there!
3.2 The Cambridge self-help guide to hustling GPUs
After two years of hustling time on Dawn, Vultr, Isambard-AI and Zenith, we've now used up the last of our Zenith preview access and are out of GPUs again. I took the spare time to write up an experience report of the operational HPC lessons learnt training Tessera, as I think we might be one of the few groups in the UK who have done large-scale training on such diverse hardware clusters. The online reaction has been positive from the community, so I've been discussing with Michael Dales where we can publish this to reach out to the wider RSE community to share learnings about academic model training.
I did upload this paper to arXiv, but it's been 'on hold' for a week. I'm not sure if that's because it's not research-y enough, or something else, but if it stays stuck there for too long I'll put this on Cambridge Open Engage or one of the other preprint archives so it has a DOI.
4 Continuing the storm of Scrutineer fixes in OCaml
After the triage in Scrutineer I started last week, lots of upstream fixes from Scrutineer are starting to flow thanks to Hannes Mehnert and Edvin Torok who are both helping now, which I'm very grateful for. Twenty-four (!) are merged, showing that the Scrutineer approach is pretty solid in terms of weeding out AI slop reports.
They mostly cluster into a few classes of problem: stuff like protocol and certificate validation logic in ocaml-tls and ocaml-x509 (crashes on malformed input, an ALPN downgrade, missing chain-depth and iteration-count limits); crypto misuse and integer/buffer bugs in mirage-crypto and digestif, losing obsolete protocols like ARC4 entirely; concurrency races and a logging deadlock in miou; memory safety in Solo5's hypervisor tooling (still open); and assorted bounds and validation fixes in ocaml-dns (open), kdf, ca-certs-nss and happy-eyeballs.
Scrutineer is gaining adoption within the OCaml community, with over a dozen developers now helping to triage issues using my private instance. I spoke to the Patch the Planet team at Trail of Bits this week, and they will pair security engineers with us in an "OCaml patching week" for the week of 19th October. If you'd like to join in, please get in touch, as we need all the help we can get right now!
Andrew Nesbitt who is the author of Scrutineer also gave me many pointers to wider ecosystem work he's doing. He has a CWE reference and a brief tool to summarise any language's build structure. Then he's got an ecosyste-ms/oss-taxonomy that could help augment my own thicket tool with a feed from live.ecosyste.ms rather than my horribly unreliable GitHub feed.
5 PR of the week
I enjoyed tinkering with the Dash for Life and sent Shane Weisz a near-me "zoom to my location" button about finding threatened species near you. Give it a try for yourself!
Next week I'll be in Oxford for a couple of days giving a talk at the Intelligent Earth CDT and then back in Cambridge!



