PERSPECTIVE article

Front. Ecol. Evol., 15 July 2026

Sec. Models in Ecology and Evolution

Volume 14 - 2026 | https://doi.org/10.3389/fevo.2026.1894720

Paleo-grounded biodiversity foundation models for long-horizon species distribution forecasting

  • 1. Polar Terrestrial Environmental Systems, Alfred Wegener Institute Helmholtz Centre for Polar and Marine Research, Potsdam, Germany

  • 2. University of Potsdam, Potsdam, Germany

  • 3. Helmholtz Centre Hereon, Geesthacht, Germany

  • 4. Helmholtz AI, Neuherberg, Germany

  • 5. Institute of Neuroscience and Medicine (INM-4), Forschungszentrum Jülich, Jülich, Germany

  • 6. Department of Psychiatry, Psychotherapy and Psychosomatics, Rheinisch-Westfälische Technische Hochschule (RWTH) Aachen University, Aachen, Germany

  • 7. Jülich Aachen Research Alliance (JARA) – CSD – Center for Simulation and Data Science, Aachen, Germany

  • 8. Technische Universität Berlin, Berlin, Germany

  • 9. BIFOLD - Berlin Institute for the Foundations of Learning and Data, Berlin, Germany

  • 10. Helmholtz Centre for Environmental Research-Umweltforschungszentrum (UFZ), Halle, Germany

  • 11. Computational Integrative Biodiversity, Goethe University Frankfurt, Frankfurt am Main, Germany

  • 12. Senckenberg Biodiversity and Climate Research Centre, Frankfurt am Main, Germany

  • 13. Landscape Ecology and Site Evaluation, University of Rostock, Rostock, Germany

  • 14. Department of Integrative Biology, University of South Florida, Tampa, FL, United States

  • 15. Institute for Biology, Martin Luther University Halle-Wittenberg, Halle, Germany

  • 16. German Centre for Integrative Biodiversity Research (iDiv) Halle-Jena-Leipzig, Leipzig, Germany

  • 17. Nederlandse Organisatie voor toegepast-natuurwetenschappelijk onderzoek (TNO), Den Haag, Netherlands

  • 18. Departamento de Botánica, Universidad de Córdoba, Córdoba, Spain

  • 19. Faculty of Geoinformation Science and Earth Observation (ITC), University of Twente, Enschede, Netherlands

  • 20. College of Life and Environmental Sciences, University of Birmingham, Birmingham, United Kingdom

  • 21. Department of Plant Biodiversity, University of Bonn, Bonn, Germany

  • 22. Center for Ecological Dynamics in a Novel Biosphere (ECONOVO), Department of Biology, Aarhus University, Aarhus C, Denmark

  • 23. Experimental Plant Ecology, University of Greifswald, Greifswald, Germany

  • 24. Helmholtz Centre for Environmental Research-Umweltforschungszentrum (UFZ), Leipzig, Germany

Abstract

Global biodiversity change demands decision-ready forecasts that integrate heterogeneous biological and environmental data across space and time. Current species distribution models (SDMs) support conservation tasks, but most remain trained and evaluated as static snapshots, whereas restoration trajectories, range shifts, and ecological baselines are shaped by dispersal limitation, disturbance, demographic inertia, and historical legacies. We argue that the next step is a paleo-grounded biodiversity foundation model: a multimodal model that learns transferable species-environment representations from modern occurrences, remote sensing, land cover, climate, and topography, and then uses paleoecological archives to test and adapt these representations across centuries to millennia. We outline two complementary temporal-transfer strategies. A conservative route trains a modern species-environment encoder and applies it to paleoclimate and paleo-land-cover slices, using pollen and sedimentary ancient DNA either for independent validation or for lightweight proxy-aware adaptation. A more ambitious route learns latent ecological dynamics from irregular, uncertainty-dated paleo time series and rolls these dynamics out across gridded landscapes. This perspective on future SDMs emphasizes proxy observation models, disequilibrium-aware evaluation, transparent uncertainty, and decision-facing outputs. By linking AI-based representation learning with paleo archives, such models could move SDMs from static suitability maps toward time-stamped, testable trajectories for biodiversity synthesis, conservation planning, and restoration monitoring.

1 Introduction

Recent biodiversity syntheses show rapid terrestrial biodiversity change and substantial uncertainty about future trajectories (). Conservation and restoration therefore increasingly depend on predictive tools that can synthesize occurrences, remote sensing, climate data, and ecological knowledge into decision-ready forecasts (). Species distribution models (SDMs) are central to reserve design, refugia mapping, invasive-risk screening, reintroduction planning, and monitoring design, but many operational workflows still produce snapshot suitability maps rather than plausible pathways through time. This is a limitation when managers need to ask when restored habitat may become colonizable, how fast a treeline may advance, or which monitoring sites can reveal whether a forecast is wrong.

The missing dimension is not only time as a coordinate, but time as an ecological process. Species distributions often lag behind climate because dispersal, demographic turnover, soil development, disturbance regimes, and biotic interactions create disequilibrium (). Pollen-based reconstructions and modeling syntheses show that postglacial forest establishment can be delayed by dispersal constraints (). Present-day vegetation can retain signatures of past climates (), and sedimentary ancient DNA indicates that northern ecosystem assembly after deglaciation took millennia (). Non-equilibrium SDMs are beginning to address such lags, but they are not yet routinely coupled to scalable, multi-species forecasting systems that can be evaluated under regime shifts ().

Paleoecological archives provide the empirical depth that modern monitoring lacks. Standardized biodiversity monitoring rarely spans more than decades, while pollen, macrofossil, and sedimentary ancient DNA (sedaDNA) records can provide site-based evidence of community change over millennia (Williams et al., 2018). These archives are valuable for testing disequilibrium and no-analog states, yet they remain underused in distribution forecasting at scale (; ; ). A perspective for AI in ecological synthesis is therefore to treat paleo archives not as historical illustrations, but as temporal supervision, stress tests, and calibration evidence for long-horizon biodiversity models.

2 From deep SDMs to biodiversity foundation models

Classical correlative SDMs often fit separate models per species with manually generated predictors. More recently, deep learning approaches have begun to learn environmental representations directly from spatiotemporal climate data, reducing the need for manually engineered predictor sets and enabling the extraction of complex ecological signals from raw environmental time series (; ). Deep multi-species approaches learn shared representations and species embeddings, allowing information to transfer across taxa with uneven data density (). Convolutional SDMs extend these point-based covariates to landscape tensors and can even capture spatial neighborhood structure from environmental raster stacks (). Remote-sensing SDMs further use satellite image time series as habitat representations, as shown for global orchid modeling with Sentinel-2 data ().These developments already move SDMs away from isolated single-species regressions toward joint learning across taxa and landscapes.

Recent models move closer to foundation-style learning. MiTREE uses a multi-input transformer encoder to combine satellite imagery, soil grids, bioclimatic variables, and ecoregion information without forcing all inputs onto one upsampled grid (). NicheFlow frames SDM as generative niche learning in environmental space, aiming to improve few- and zero-shot prediction for poorly sampled species (). In parallel, geospatial foundation models show that large transformer backbones can be pretrained on massive Earth-observation corpora and transferred to downstream tasks (). Presto () illustrates efficient pixel-level remote-sensing time-series modeling with structural masking for missing sensors and timesteps, while the broader geo-foundation model literature emphasizes explicit representations of space, time, and transfer ().

The closest current step toward a biodiversity foundation model is BioAnalyst, which aligns multimodal environmental data with large-scale occurrence supervision and supports lightweight downstream adaptation (). We use the term biodiversity here to represent the modeling goal of entire species compositions, which can be used to derive community dynamics or biodiversity indices. Yet, BioAnalyst’s horizon remains primarily modern. The key gap is not only multimodal pretraining, but temporal transfer across observation regimes: modern presence-only occurrences, paleo proxy counts, sedaDNA reads, paleoclimate simulations, and reconstructed land cover do not measure the same ecological object in the same way (). General-purpose chatbots cannot fill this gap, because they are not trained as spatially explicit ecological models and can produce overconfident or biased conservation narratives ().

3 Data streams for a paleo-grounded model

A practical model should begin with a modern species-environment encoder that learns from gridded climate, topography, land cover, and Earth-observation context, while being explicit about spatial grain and missing inputs. A multitude of modern environmental variables are already available globally and at high spatial resolutions. CHELSA-BIOCLIM+ provides global climate-related predictors at kilometer resolution (), and modern annual land-cover classifications from ESA CCI/Copernicus C3S can represent habitat and land-use structure rather than climate alone (). Topographic conditioning remains important for both modern habitat structure and paleo downscaling, while glacial-isostatic reconstructions can provide long-term boundary conditions where terrain and coastlines changed substantially (). Modern species occurrence data sets also provide broad coverage but encode sampling effort, detectability, taxonomic curation, and accessibility biases at the same time; these biases must be included in the model rather than treated as an ecological signal.

The model’s paleo extension must account for data that are not simply older versions of modern occurrences. Paleo proxy evidence is characterized by varying taxonomic, spatial and temporal resolution as well as substantial dating uncertainty. Pollen records capture local-to-regional vegetation signals through counts or relative abundances, whereas sedimentary ancient DNA (sedaDNA) can complement pollen by providing different taxonomic resolution and detection processes (). The Neotoma Paleoecology Database provides the primary harmonized access layer for paleo records across proxies, supporting standardized metadata and chronologies (Williams et al., 2018).

Modeling long-term ecological dynamics requires covariates that are consistently available across millennial timescales (). Paleoclimate surfaces such as PaleoClim and CHELSA-TraCE21k provide temporally explicit climate drivers for the late Quaternary (; ). Because paleo evidence is spatially and taxonomically uneven, trait information and phylogenetic structure can provide important inductive bias for shared environmental tolerances and response times (; ; ).

The explicit inclusion of land cover within the temporal predictor stack may further improve model robustness. This could incorporate: (a) modern annual land cover classifications derived from Copernicus C3S/ESA CCI products, and (b) pollen-based paleo land cover reconstructions generated through statistical reconstruction and spatial interpolation using U-Net models (Schild et al. in prep). Together, these datasets provide additional spatially and temporally consistent information relevant to species occurrence dynamics.

Our central recommendation is to model these data streams as linked but non-equivalent observations. A modern occurrence point, a pollen count, and a sedaDNA read are all evidence about past or present biotic states, but each has its own detection process, spatial footprint, taxonomic resolution, and dating uncertainty (see Figure 1). Treating them as clean presence-absence labels would make the model look precise while encoding false confidence. The foundation-model objective should therefore include proxy-aware likelihoods, masks for missing modalities, explicit age distributions for paleo samples, and taxonomic reconciliation layers that can connect genus-level proxy evidence to species-level modern labels without inventing unsupported precision.

Figure 1

(licensed under CC-BY-4.0); (CC0 1.0); (GNU General Public License V3); (CC0 1.0); (licensed under CC-BY-4.0); and (licensed under CC-BY-4.0).

A second data principle is to keep ecological and sampling uncertainty separable. Presence-only observations may reflect collection effort as much as species abundance, whereas pollen percentages may reflect production, transport, preservation, and counting depth. A useful model should therefore represent observation effort where possible, learn taxon-specific detectability parameters where evidence allows, and expose low-confidence regions rather than smoothing them away. This is especially important for underrepresented regions and taxa, where a foundation model can otherwise appear more general than the data justify.

4 A model architecture for temporal transfer

We propose a modular architecture built around a latent ecological state (see Figure 2). The latent state provides a compact way to encode interacting ecological information, such as climate, land cover, community composition, disturbance, and legacy effects, without requiring all processes to be specified explicitly.

Figure 2

The first module encodes modern environmental context and occurrence supervision into a transferable representation of habitat and community conditions. A U-Net provides a robust spatial baseline because encoder-decoder convolutional networks can combine broad context with fine spatial detail for dense prediction (). This is useful for ecological data because species distributions depend not only on local conditions, but also on surrounding landscape context, including connectivity, fragmentation and refugia. A 3D U-Net can represent short regular space-time blocks, but long paleo sequences and irregular samples make this expensive and restrictive (). A useful implementation can therefore begin with a U-Net baseline to verify the data pipeline, then replace or augment the encoder with a multimodal transformer when the evaluation shows that flexible modality fusion is needed.

A transformer-based implementation is better suited to heterogeneous and partially missing modalities. This flexibility is important for ecological, and especially paleoecological, data, where observations differ in spatial coverage, temporal resolution, uncertainty and modality. Attention mechanisms support flexible fusion across input types (), and vision transformers can learn patch-based spatial structure when supported by sufficient pretraining (). Masked autoencoding and data-efficient transformer training are useful because ecological data are incomplete by design (; ). Hierarchical transformers reduce the cost of large raster inputs (), while Perceiver IO compresses arbitrary multimodal inputs into a fixed latent array that can feed task-specific outputs (). The architecture should keep the species head compact, for example by predicting shared community maps or latent composition factors that are projected to species probabilities, because full per-species maps for thousands of taxa are costly and hard to calibrate.

Training should combine self-supervision and ecological supervision. Self-supervised objectives can reconstruct masked environmental layers, missing timesteps, or withheld modalities, which teaches the encoder spatial and temporal context before species labels dominate the loss. Supervised objectives can then connect the latent state to modern occurrences, while bias terms or background-sampling strategies reduce the risk that the model learns where people observe rather than where species can occur. Paleo losses should be added later and sparingly, because proxy data are powerful for temporal validation but too heterogeneous to be treated as a simple extension of modern occurrence data.

The temporal component can then be developed through two complementary routes. Temporal learning is needed because species distributions are rarely in equilibrium with current environments; they also reflect dispersal limitation, demographic inertia, disturbance history, and ecological memory. Strategy A is modern-first hindcasting. The modern encoder is trained on contemporary occurrence-environment links, then applied to paleoclimate and paleo-land-cover slices. Paleo proxy data are withheld for independent tests, or used only through lightweight adapters and proxy-aware output layers that align predictions to pollen or sedaDNA observation processes. This route is conservative because it asks whether modern-learned associations survive transfer to past climates, and it makes failure informative rather than hiding it inside a larger end-to-end model. It also guards against catastrophic forgetting: the model can be adapted to proxy evidence without overwriting modern predictive skill.

Strategy B learns latent dynamics from site time series. Each paleo site provides irregular, uncertainty-dated samples; for each sample, the encoder produces a latent ecological state conditioned on the available drivers. A temporal transition model then learns how state(t) changes to state(t+dt) under changing climate and land cover. This transition could be a state-space model, an irregular-time transformer, or a continuous-time recurrent module, but it should be trained against proxy likelihoods rather than direct labels. Once learned at sites, the transition can be rolled out across a spatial grid to generate coherent three-dimensional distribution fields through time. This route is more ambitious because it treats paleo archives as evidence for a rule of ecological change, not merely as validation points.

These strategies are complementary rather than competing. Strategy A gives an interpretable baseline for how far modern niche relationships can be pushed, while Strategy B tests whether the model can learn path dependence and response lags. The strongest perspective is to use both: if Strategy A fails in a region where Strategy B succeeds, the difference points to temporal dynamics rather than only environmental mismatch; if both fail, the system should flag a domain where the model is not reliable.

5 Evaluation, outputs, and applications

Evaluation should be designed around transfer, calibration, and usefulness. Modern performance should be assessed with spatial and temporal hold-outs, independent detections where available, and bias-aware discrimination and calibration metrics. Paleo performance should be assessed by hindcasting to dated pollen and sedaDNA time series while withholding sites, regions, or time intervals. Baselines should be classical SDMs fitted with matched environmental drivers, and ablations should test whether proxies, traits, phylogeny, and land-cover history improve transfer rather than only in-sample fit. Posterior predictive checks should ask whether simulated proxy observations resemble real proxy records, not only whether latent suitability maps look plausible.

A useful benchmark should therefore be organized as a ladder. The first rung tests modern spatial transfer; the second tests modern temporal transfer; the third tests paleo hindcasts with proxies withheld; the fourth tests proxy-aware adaptation on one set of sites and prediction to another; and the fifth tests taxonomic transfer to sparsely observed taxa. Each rung should report discrimination, calibration, trajectory coherence, and uncertainty coverage. The goal is not to declare a single model winner, but to identify when a foundation-model representation is actually transferable and when classical SDMs remain more reliable.

Decision-facing outputs should separate habitat suitability, accessibility, and uncertainty. Time lags in SDM calibration can mislead conservation decisions when realized ranges trail environmental change (). A paleo-grounded model should therefore report where habitat appears suitable, where colonization is plausible given inferred lags and barriers, and where uncertainty is high enough to justify monitoring. In floodplain wetland restoration, such outputs could help distinguish habitat limitation from dispersal limitation along hydrological gradients. Under Arctic treeline advance, they could map probabilistic shrub and tree encroachment fronts, tundra refugia, and the expected lag between climatic suitability and realized vegetation change.

The model must also be usable outside the modeling team. Outputs should be exported as standard geospatial products, such as Cloud Optimized GeoTIFF maps for occurrence probability and uncertainty and NetCDF time series for decadal or centennial trajectories (). Interactive viewers and natural-language helpers can lower the barrier for practitioners, but the helper should translate questions into model runs and visual summaries rather than generating unsupported ecological claims. This distinction is essential because large language interfaces can make uncertainty look conversationally settled even when the underlying prediction is weak.

For applied users, the most valuable product may not be a single forecast but a comparison of plausible trajectories. A restoration manager may need to compare passive recovery, assisted colonization, and hydrological intervention; a conservation planner may need to compare climate scenarios and identify monitoring sites that most reduce uncertainty. A paleo-grounded model is well suited to this use because hindcasts provide a way to test whether the same workflow could have reconstructed known historical transitions before it is trusted for future ones.

6 Discussion

The main feasibility challenge is not whether AI can ingest more data, but whether ecological meaning is preserved while data streams are fused. Three safeguards are essential. First, observation models must remain proxy-specific so that pollen, sedaDNA, and modern occurrences are not collapsed into one artificial label space. Second, evaluation must include spatial, temporal, taxonomic, and proxy hold-outs so that transfer is measured under realistic extrapolation. Third, uncertainty should be communicated as part of the output rather than as an appendix.

The near-term research agenda should focus on shared benchmarks rather than only larger models. The community needs harmonized modern and paleo data splits, agreed proxy observation models, transparent rules for taxonomic reconciliation, and baseline comparisons under matched drivers. It also needs negative results: cases where modern niche representations fail in the past are scientifically useful because they reveal disequilibrium, missing drivers, or no-analog communities. Interdisciplinary collaboration is not optional here, because machine learning expertise alone cannot define valid proxy likelihoods, and paleoecological expertise alone cannot scale multimodal representation learning across taxa and continents.

A paleo-grounded biodiversity foundation model would not replace mechanistic ecology or local expertise. It would provide a synthesis layer that turns scattered modern and paleo observations into testable spatiotemporal hypotheses. Unlike traditional SDMs, which usually produce static suitability maps from contemporary species–environment correlations, this framework would model species distributions as uncertainty-aware trajectories across past, present, and future conditions. By interpolating distributions through time, it could account for ecological lags, historical baselines, and range-shift dynamics, supporting both forward-looking conservation planning and retrospective analyses of past biodiversity change. Ethical deployment requires decision-appropriate spatial resolution, protection of sensitive species localities, explicit documentation of underrepresented regions and taxa, and collaboration with local and Indigenous knowledge holders. Used in this way, AI can help conservation move beyond static suitability maps toward historically informed, uncertainty-aware trajectories that are scientifically testable and practically actionable.

Statements

Data availability statement

The original contributions presented in the study are included in the article/Supplementary Material. Further inquiries can be directed to the corresponding author.

Author contributions

UH: Visualization, Data curation, Writing – review & editing, Conceptualization, Investigation, Writing – original draft, Supervision. LS: Data curation, Methodology, Visualization, Conceptualization, Investigation, Writing – review & editing, Writing – original draft. LF: Data curation, Writing – review & editing, Conceptualization, Investigation. DC: Writing – review & editing, Conceptualization, Methodology. JX: Data curation, Conceptualization, Writing – review & editing. TW: Writing – review & editing, Methodology, Conceptualization. JD: Writing – review & editing, Conceptualization. BD: Conceptualization, Writing – review & editing. MG: Conceptualization, Writing – review & editing. VG: Writing – review & editing, Conceptualization. OH: Conceptualization, Writing – review & editing. FJ: Writing – review & editing, Conceptualization. TK: Conceptualization, Writing – review & editing. AK: Conceptualization, Writing – review & editing. IK: Writing – review & editing, Conceptualization. MM: Conceptualization, Writing – review & editing. DN-L: Writing – review & editing, Conceptualization. CP: Writing – review & editing, Conceptualization. BR: Conceptualization, Writing – review & editing. JC: Writing – review & editing, Conceptualization. BS: Writing – review & editing, Conceptualization. J-CS: Conceptualization, Writing – review & editing. FTan: Writing – review & editing, Conceptualization. FTau: Conceptualization, Writing – review & editing. DZ: Writing – review & editing, Conceptualization.

Funding

The author(s) declared that financial support was received for this work and/or its publication. This work was supported by Helmholtz Association’s Initiative and Networking Fund through Helmholtz AI [grant number: ZT-I-PF-5-01], the Federal Ministry of Research, Technology and Space 373 (BMFTR 03F0944A), and the PalMod Initiative (grant no. 01LP1510C). J-CS was supported by Center for Ecological Dynamics in a Novel Biosphere (ECONOVO), funded by Danish National Research Foundation (grant DNRF173).

Acknowledgments

The authors acknowledge support by the Open Access publication fund of Alfred-Wegener-Institut Helmholtz-Zentrum für Polar- und Meeresforschung.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

The author JD declared that they were an editorial board member of Frontiers at the time of submission. This had no impact on the peer review process and the final decision.

Generative AI statement

The author(s) declared that generative AI was used in the creation of this manuscript. text drafting and editing.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fevo.2026.1894720/full#supplementary-material

References

Summary

Keywords

biodiversity change, conservation planning, foundation models, paleoecology, species distribution models

Citation

Herzschuh U, Schild L, Farkas L, Caus D, Xia J, Weigel T, Dammers J, Demir B, Golivets M, Gupta V, Hagen O, Jansen F, Khan T, Kramer A, Kühn I, Mensio M, Nieto-Lugilde D, Persello C, Rasti B, Cabral JS, Schröder B, Svenning J-C, Tanneberger F, Taubert F and Zurell D (2026) Paleo-grounded biodiversity foundation models for long-horizon species distribution forecasting. Front. Ecol. Evol. 14:1894720. doi: 10.3389/fevo.2026.1894720

Received

29 May 2026

Revised

23 June 2026

Accepted

25 June 2026

Published

15 July 2026

Volume

14 - 2026

Edited by

Nataša Popović, Institute for Biological Research “Siniša Stanković”, Serbia

Reviewed by

Anna Maria Mercuri, University of Modena and Reggio Emilia, Italy

Updates

Copyright

*Correspondence: Ulrike Herzschuh,

†These authors have contributed equally to this work

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics