Michael Dales, Aneesh Naik and Alison Eyres are working with David Coomes' group and me on whether the TESSERA foundation model can produce robust, high-resolution (10m) global habitat maps.
Habitat maps underpin most of the conservation pipelines we work on (such as Mapping LIFE on Earth and its area-of-habitat calculations), but the maps currently available are either coarse, dated, or stitched together from sources with varying definitions of what a "habitat" actually is. A pixelwise embedding of the planet's land surface at 10m resolution gives us a chance to build them more directly from observations, and without relying too much on spatial context.
The pipeline we are prototyping takes species occurrence records from GBIF, filters them against IUCN Red List range polygons and habitat preferences, and uses the survivors as (weak and noisy) labels over Tessera embeddings to classify habitat.
Early experiments surface many aspects of having to clean up citizen-science data. For example, over a million UK occurrences for 182 species collapse to ~58,000 once range-filtered, and those cluster heavily around nature reserves, canals and even Chester Zoo (spatial bias around where observations happen!). We're making these biases visible early in the pipeline, rather than discovering them after aggregation, via a data cleanup pipeline.
We are starting with UK maps, where we have good intuitions about what the answers should look like. Madagascar is next, working with Julia P.G. Jones and Miranda Lam, as a contrasting region where better habitat maps would have a direct conservation impact.
Michael writes up progress in his weeknotes, as is Aneesh on his site.
1 See also
James G. C. Ball's work on data-efficient tree species mapping is essential background reading for building classifiers on TESSERA embeddings, and the broader TESSERA programme this builds on.
