SYSTEMATIC REVIEW article

Front. Clim., 21 July 2026

Sec. Climate and Health

Volume 8 - 2026 | https://doi.org/10.3389/fclim.2026.1843933

Environmental predictors and machine learning models for climate-health risk prediction from extreme weather events: a global systematic review

  • 1. School of Computing, Engineering and Physical Sciences, University of the West of Scotland, Paisley, United Kingdom

  • 2. School of Agriculture and Science, University of KwaZulu-Natal, Durban, South Africa

  • 3. School of Electrical and Mechanical Engineering, University of Portsmouth, Portsmouth, United Kingdom

  • 4. School of Health and Life Sciences, University of the West of Scotland, Paisley, United Kingdom

  • 5. Faculty of Nursing and Midwifery, Royal College of Surgeons, University of Medicine and Health Sciences, Dublin, Ireland

  • 6. Department of Civil and Environmental Engineering, University of Strathclyde, Glasgow, United Kingdom

  • 7. Discipline of Public Health, School of Medicine, University of KwaZulu-Natal, Durban, South Africa

Abstract

Climate change is increasing the frequency, severity, and duration of extreme weather events, including heatwaves, floods, and heavy rainfall, posing substantial risks to population health worldwide. Traditional epidemiological approaches capture retrospective associations but are limited in modeling the non-linear, multivariable relationships between climate exposures and health outcomes. This review synthesized evidence on the application of machine learning models to predict health outcomes associated with extreme weather events and identify the environmental predictors most consistently reported as important. A comprehensive search across 11 electronic databases was conducted following the PRISMA-2020 guidelines, identifying peer-reviewed studies published between July 2010 and July 2025. Eighteen studies met the inclusion criteria and were assessed for risk of bias using the Prediction Model Risk of Bias Assessment Tool. Heat-related exposures dominated the included studies, with 17 of the 18 studies focusing on heatwaves, while only one examined heavy rainfall in relation to dengue outcomes. Temperature-based variables were consistently used as predictors across all studies. Socioeconomic and demographic variables were included in 10 studies, but were ranked as the most influential predictors in only three. Random Forest was the most frequently evaluated algorithm and was identified as the best-performing model in seven studies spanning high-income and upper-middle-income settings. However, the Generalized Additive Model outperformed machine learning models in two studies, and predictive performance varied substantially across contexts. Although the evidence base is concentrated in high-income countries, external validation was absent across all included studies, limiting the confidence in model generalizability. Despite the dominance of heat-related studies, the absence of models for flood-related health outcomes represents a critical gap, given the global burden of flood-related morbidity. Key limitations include limited integration of socioeconomic determinants, uneven geographical distribution of evidence, and sparse data availability with coarse spatial resolution in low- and middle-income countries, constraining model transferability. Significant variability in predictive performance, combined with the absence of external validation, indicates that the operational deployment of these models remains premature. Addressing these gaps requires expanding the modeling approach to include rainfall and flood-related health outcomes, improving data availability and spatial resolution, and developing transferable model architectures to support reliable early warning systems.

Systematic review registration:

PROSPERO https://www.crd.york.ac.uk/PROSPERO/view/CRD420251077655, identifier CRD420251077655.

1 Introduction

Climate change has emerged as one of the most significant global public health challenges of the 21st century (Weeda et al., 2024; Khatibu and Ngowi, 2025; ). The Intergovernmental Panel on Climate Change (IPCC) reports that human-induced climate change is increasing the frequency, intensity, and duration of Extreme Weather Events (EWEs) worldwide, including heatwaves, floods, and heavy rainfall (Intergovernmental Panel on Climate Change (IPCC), 2023; Kafi and Ponrahono, 2024; Seneviratne et al., 2021). These events are associated with a range of acute health outcomes directly attributable to specific EWE episodes, including heat-related illnesses, respiratory complications, infectious disease transmission, injuries, and increased mortality (Rocha et al., 2022; Yenew et al., 2025; Franchini and Mannucci, 2015; Xi et al., 2024). Vulnerable populations, including the elderly, children, and socioeconomically disadvantaged communities, face disproportionate risks owing to their limited adaptive capacity and increased exposure (Kirby et al., 2025; World Health Organization, 2024; Dickinson et al., 2025). However, anticipating when and where these health impacts will occur remains a major challenge, highlighting the need for predictive approaches capable of identifying climate-sensitive health outcomes before they emerge.

Conventional epidemiological studies have largely examined associations between environmental exposures and observed health outcomes using retrospective analytical frameworks such as time-series, case-crossover, and related statistical approaches (Berman et al., 2017; Lawrence et al., 2021). Environmental predictors serve as fundamental input variables for climate-health predictive models (), yet their identification, selection, and relative importance across different health outcomes and geographic contexts remain poorly understood. Studies have incorporated a wide range of environmental input variables, including meteorological parameters such as temperature, humidity, and precipitation as well as atmospheric and air quality variables such as air quality indicators, solar radiation, and atmospheric pressure (), but they often lack systematic assessment of their individual and combined predictive value across modeling approaches (Camps-Valls et al., 2025). Furthermore, the availability, quality, and resolution of environmental and health surveillance data differ markedly between high-income countries (HICs) and low- and middle-income countries (LMICs) (Ciecierski-Holmes et al., 2022; Kaushik et al., 2025), reflecting disparities in meteorological station density, satellite data access, and health monitoring infrastructure, which collectively constrain modeling capacity and weaken global health preparedness (; Bartlow et al., 2019).

To address these challenges, machine learning (ML) and deep learning (DL) offer considerable potential for predicting extreme weather-related health outcomes by capturing complex non-linear relationships that traditional techniques often overlook (Ssebyala et al., 2024; Nusrat et al., 2022). ML and DL models can process large streams of continuously updated environmental data in real time, supporting the development of early warning system (EWS) designed to protect vulnerable populations during EWEs (Darsha Jayamini et al., 2024; Liu et al., 2025; El Morr et al., 2024). Recent advances in ensemble modeling, feature selection, and model interpretation methods, such as SHapley Additive exPlanations (SHAP) (Im et al., 2025), and permutation importance (Boudreault et al., 2024), have demonstrated strong potential for improving prediction performance and transparency in climate-health modeling across diverse geographic and socioeconomic settings (Teshale et al., 2024; Boudreault et al., 2023). Within the climate-health domain, ML has been applied to predict EWEs that precede health impacts, such as forecasting long-lasting extreme heatwaves up to 15 days in advance using a neural network (NN) from surface temperature (Jacques-Dumas et al., 2022), and predicting flood-driven streamflow and community mobility disruption from satellite remote sensing and rainfall data in rural Rwanda (Macharia et al., 2023). ML has also been applied to model health outcomes directly, including heatwave-related mortality using climate variability and oscillation indices (Boudreault et al., 2023). While EWE prediction provides important inputs for health risk modeling, this review focuses specifically on studies that directly predict health outcomes. These developments highlight the growing role of ML not only as a predictive tool but also as a framework for developing interpretable, scalable, and potentially deployable health outcome prediction models (Ciecierski-Holmes et al., 2022).

Despite this growing methodological literature, the evidence base remains fragmented across geographic coverage, outcome definitions, modeling approaches, and predictor selection. Prior reviews have examined infectious disease surveillance without considering weather-related exposures (El Morr et al., 2024; Villanueva-Miranda et al., 2025), occupational heat stress in workers rather than population-level extreme weather risks (Ferrari et al., 2025), or broad climate-health domains without extracting environmental predictors or evaluating ML model performance (Berrang-Ford et al., 2021). Most recently, a comprehensive review synthesized environmental predictors and compared algorithm performance across 25 heat-focused studies, but was restricted to heat exposure and lacked formal methodological quality appraisal (Boudreault et al., 2025). A scoping review on ML methods for EWE-associated health risks identified only seven eligible studies, reporting health outcomes, but neither mapped environmental predictors to specific health outcomes nor assessed methodological quality (Ssebyala et al., 2024). Critically, no prior review has applied a validated risk of bias assessment tool to assess the methodological quality of EWE-health prediction models, systematically characterized the structural barriers preventing flood-health prediction modeling, including data availability constraints, and health surveillance limitations or evaluated cross-contextual model generalizability in terms of geographic transferability and socioeconomic settings. Consequently, fundamental questions remain unanswered: which environmental predictors are most influential for EWE-associated health outcomes, how do different ML algorithms compare in predictive performance across contexts, and which factors constrain model transferability and adaptation to LMIC contexts?

This systematic review addresses these gaps by synthesizing evidence on ML models used to predict specific health outcomes associated with EWEs, including mortality, morbidity, and healthcare utilization, with particular emphasis on the environmental predictors driving model performance and cross-contextual generalizability of the models. Specifically, this review aims to: (1) identify and categorize ML models employed in predicting EWE-associated health outcomes; (2) determine the most commonly used environmental predictors, including meteorological variables, air quality indicators, and satellite-derived variables; (3) identify the most influential predictors based on feature importance methods reported in primary studies; (4) examine how environmental predictors from specific EWEs map to modeled health outcomes; and (5) evaluate model predictive performance using reported metrics and assess generalizability across diverse geographic and socioeconomic settings. The findings aim to support the development of context-appropriate predictive models and provide evidence to inform the design of effective EWS for climate-sensitive health risks worldwide.

2 Methods

The review process was organized into three phases: planning, conducting, and reporting, following an established systematic review methodology framework (Kitchenham et al., 2004), as illustrated in Figure 1. The planning phase involved defining the review objectives and eligibility criteria. The conducting phase included a database search, study selection, data extraction, and quality assessment. The reporting phase focused on synthesizing the findings and preparing the final manuscript. The methodological components are described in detail in the following subsections.

Figure 1

2.1 Protocol registration and reporting standards

The protocol for this systematic review was prospectively registered in PROSPERO (CRD420251077655) (), prior to study selection, and no amendments were made to the registered protocol. The review followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyzes (PRISMA-2020) guidelines for overall reporting (Page et al., 2021).

2.2 Information sources and search strategy

The PECO (Population, Exposure, Comparator, and Outcome) framework was used to refine the research question and inform the development of search terms (Morgan et al., 2018). The PECO framework was selected in preference to PICO (Population, Interventions, Comparator, and Outcomes) because this review examines observational predictive modeling studies in which no active intervention or direct comparator was applicable; therefore, PECO is more appropriate than PICO for structuring evidence synthesis in environmental exposure research (Morgan et al., 2018). A comprehensive literature search was initiated on 1 July 2025, beginning with the Scopus database. An initial set of search terms was developed and reviewed by a specialist librarian using PRESS-2015 guidelines (McGowan et al., 2016), to ensure accuracy and completeness. Following refinement, the full search strategy was finalized and implemented across databases on 3 July 2025. We searched eleven databases spanning biomedical, environmental and multidisciplinary literature: Scopus, Web of Science, IEEE Xplore, PubMed, ScienceDirect, Wiley Online Library, MEDLINE, GreenFILE, CINAHL Ultimate, Education Source, and LISTA, covering the period from July 2010 to July 2025. Education Source and LISTA were included to capture interdisciplinary studies on environmental exposure and modeling approaches not consistently indexed in biomedical databases. The search was limited to studies published from 2010 onwards, as this period coincides with the acceleration of ML applications in healthcare research and the emergence of DL techniques following key algorithmic breakthroughs between 2010 and 2012 (LeCun et al., 2015; Jordan and Mitchell, 2015).

The search terms covered environmental predictors, exposure to EWEs, such as extreme rainfall, flooding, and heatwaves (including indoor heat exposure resulting from outdoor conditions), ML methods, and human health outcomes. Boolean and proximity operators were used to structure the search syntax. An example of the Scopus search strategy is presented in Table 1. Searches were limited to English peer-reviewed journal articles, and citation exports were filtered at the database level where possible (e.g., document type, restricted to subject areas). All eligible records were exported in .RIS format to facilitate direct import into the Rayyan tool for further processing (Ouzzani et al., 2016). Full details of the database-specific search strategies and refinement procedures are provided in Supplementary file S1.

Table 1

PECOSearch componentSearch syntax (TITLE-ABS-KEY)
PExtreme weather events(“heatwave*" OR “heat wave*" OR “extreme heat" OR “high* temperature" OR “flood*" OR “heavy rain*" OR “extreme weather event*" OR “climate change" OR “weather change") AND
EEnvironmental predictors(“humidity" OR “precip*" OR “wind speed" OR “wind velocity" OR “solar radia*" OR “climate data" OR “health data" OR “environmental varia*" OR “weather data" OR “meteorological data" OR “enviro* paramet*" OR “rainfall intens*" OR “air quality" OR “atmospheric pressure") AND
EAI and ML terms(“ML" OR “deep learning" OR “ensemble learning" OR “artificial intelligence" OR “AI" OR “neural network*" OR “random forest" OR “support vector machine*" OR “KNN" OR “Naive Bayes" OR “gradient boosting" OR “predictive model*" OR “forecast*" OR “data-driven model*" OR “climate model*" OR “weather model*" OR “health risk model*" OR “decision tree" OR “classification algorithm*") AND
CComparatorNot applicable
OHealth outcomes(“health" OR “health risk*" OR “public health" OR “human health" OR “vulnerab*" OR “morbidity" OR “mortality" OR “hospita*" OR “emergency visit*" OR “symptom*" OR “illness*" OR “health impact*" OR “infect*") AND
Filters appliedLIMIT-TO (SUBJAREA,“ENVI" OR “EART" OR “ENGI" OR “MEDI" OR “SOCI" OR “COMP" OR “MATH" OR “HEAL") AND LIMIT-TO (DOCTYPE,“ar") AND LIMIT-TO (LANGUAGE,“English")

Scopus search strategy used in this review.

* denotes a wildcard (truncation) symbol used in the search strategy to retrieve multiple word variations sharing the same root.

2.3 Eligibility criteria

Table 2 summarizes the eligibility criteria applied in this review. As this review focuses on predictive modeling studies rather than comparative interventions, no comparator was applicable. The primary objective was to identify studies that applied ML methods to predict human health outcomes associated with EWEs, using environmental or meteorological predictors. Eligible studies were required to use real-world weather or environmental monitoring data during extreme weather periods and link these exposures to population-level health outcomes. No restrictions were applied on geographic region, age and gender.

Table 2

DomainInclusion criteriaExclusion criteria
Population (P)Human populations exposed to EWEs (heatwaves, floods, and extreme rainfall), with no restrictions on age, gender, or location.Studies based on infrastructure, agriculture, animal populations, or non-human subjects.
Exposure (E)Studies must use environmental or meteorological variables as ML predictors (e.g., temperature, humidity, rainfall, wind, pressure, air quality, and derived heat indices).Studies that do not include environmental predictors, or rely solely on demographic, clinical, laboratory, or indoor simulation data.
Comparator (C)Not applicable; this review examines predictive models rather than comparative interventionsNone
Outcome (O)Human health outcomes associated with extreme weather exposure, including mortality, morbidity, disease incidence, hospitalizations, and emergency events.Studies not related to human health outcomes.
ContextStudies must be conducted in real-world climatic environments where exposure is derived from population-level outdoor conditions (e.g., meteorological stations, environmental agencies).Artificial indoor experiments, controlled chamber studies, or contexts not representing actual outdoor weather or extreme events.
Publication typePeer-reviewed full-text original research reporting at least one performance metric.Reviews, conference abstracts, editorials, commentaries, case reports, and gray literature.
Time2010–2025Before 2010 were not considered
LanguageEnglish language publications onlyNon-English publications

Eligibility criteria for study inclusion and exclusion.

To be included, studies needed to employ multiple environmental parameters rather than relying solely on biomarkers, imaging, or genetic variables, and they were required to report at least one model performance metric (e.g., AUC, sensitivity, or calibration measures). Where available, both discrimination and calibration metrics were extracted. Only peer-reviewed, full-text original research articles published in English were eligible. Studies involving model validation, modification, updating, or comparison with conventional statistical approaches were included, provided performance metrics were presented.

Studies were excluded if they were non-English or non-human, consisted of reviews, conference abstracts, case reports, editorials, or other non-peer reviewed sources or used laboratory-based or indoor-only simulation environments without real-world meteorological data. Studies focused on post-disaster recovery not linked to EWEs, as well as studies using ML solely for climate prediction or environmental monitoring without a health-related outcome, were also excluded.

2.4 Study selection process

All records retrieved from the database searches were imported into the Rayyan tool for reference management and screening. Automated deduplication was performed within Rayyan using title, author, and DOI matching, followed by manual verification of flagged records, after which unique records were screened. Two reviewers (FA and AA) independently screened the titles and abstracts, applying the predefined eligibility criteria. Conflicts were resolved through discussion with a third reviewer (MZS). After title and abstract screening, eligible reports were sought for retrieval and successfully obtained for full-text assessment. The same two reviewers independently evaluated each full text against the inclusion and exclusion criteria specified in the registered protocol. Any disagreements or uncertainties were resolved by consensus of the three reviewers. Following completion of the screening process, Google Scholar and the reference lists of all included studies were searched to identify additional relevant studies that may have been missed. Records identified through this supplementary search underwent the same full-text eligibility assessment.

2.5 Data extraction and synthesis

Data were extracted from the included studies using a standardized extraction form developed a priori and piloted on three studies. Data extraction was performed by one reviewer (FA) and independently checked by a second reviewer (AA) for consistency. Any discrepancies were resolved through discussion with a third reviewer. The extraction form was implemented primarily in Excel spreadsheets, which were used to record and structure all extracted variables. Extracted information was organized into predefined domains consistent with the PECO framework and the review protocol: study identification and research scope (authors, year, country, aim, and sample size), EWE characteristics and study settings (event type, geographic scale, study design, and study period), environmental and meteorological predictors (predictor variables, spatial and temporal resolution, and derived or composite indicators), AI and ML model characteristics (model type, validation method, training and testing split, and performance metrics), and health outcomes (outcome definition, data source, population setting, and prediction time horizon), detailed in Supplementary file S4. Extracted data were maintained locally and cross-referenced with Rayyan records to ensure consistency and traceability across the screening and extraction stages.

Data were synthesized narratively following the protocol-specified approach due to substantial heterogeneity across ML model types, environmental predictors, health outcomes, and model validation methods. For studies evaluating multiple algorithms, all reported models were extracted, and the best-performing model was additionally identified and reported based on the performance metrics provided in the included studies. Studies were grouped thematically according to EWE types and health outcome to facilitate comparison of predictor selection patterns, model performance, validation strategies, and evidence of geographic transferability. Quantitative synthesis was not undertaken because no outcome grouping met the minimum homogeneity requirements defined in the protocol. Although some groupings contained three or more studies, such as all-cause mortality and heat-related mortality, these studies differed considerably in ML architectures, environmental predictors, spatial and temporal resolution, validation strategies, performance metrics, outcome definitions, and study design characteristics. As no grouping demonstrated sufficient methodological, clinical, or contextual similarity to allow meaningful statistical pooling, meta-analysis was deemed inappropriate and narrative synthesis was retained.

2.6 Quality assessment

Risk of bias and applicability were assessed using the Prediction Model Risk of Bias Assessment Tool (PROBAST) (Wolff et al., 2019). PROBAST evaluates prediction model studies across four domains: participants, predictors, outcomes, and analysis, using structured signaling questions to identify potential sources of bias. Applicability was assessed for the participants, predictors, and outcome domains in relation to the objectives of this review. Studies were classified as having low, high, or unclear risk of bias. A summary of PROBAST domain-level judgments is presented in Supplementary files S2, S3.

3 Results

The database search identified 9,019 records, of which 5,737 remained after removal of duplicates (n = 3,282). Initial screening on the basis of titles and abstracts excluded 5,650 records, with 87 articles proceeding to full-text assessment. Following a full-text review against predefined eligibility criteria, 16 studies met all inclusion requirements. Two additional studies were identified through a search of Google Scholar, resulting in a total of 18 studies included in the final review. The study selection process is summarized in Figure 2, and was conducted in accordance with the PRISMA-2020 guidelines (Page et al., 2021). Details of the excluded studies during the full-text review are provided in Supplementary file S5.

Figure 2

3.1 Study characteristics and geographic distribution

The 18 included studies span the period from 2018 to 2025, with a marked rise in publications during the past three years, as shown in Figure 3. A detailed summary of the study characteristics is provided in Table 3. The included studies analyzed health data from nine countries across multiple continents: China (Wang et al., 2019; Xu et al., 2024; Wang et al., 2025), Japan (Nishimura et al., 2021; Ohashi et al., 2023; Ke et al., 2023), the United States (Wertis et al., 2023; ), Canada (Boudreault et al., 2024, 2023; Côté et al., 2024), Germany (Schachtschneider et al., 2024; Wang et al., 2024), India (Sophia et al., 2025), South Korea (Kim and Kim, 2022), Australia (Wang et al., 2023; Jian et al., 2023), and Senegal (Toure et al., 2025). The detailed geographic distribution is shown in Figure 4.

Figure 3

Figure 4

Table 3

ReferencesStudy periodGeographic scaleEWE type and definitionHealth outcomeTarget population/vulnerability
()1987–200582 large urban communities (national)Heatwave: ≥ 2 days Tmean ≥ 98th percentileAll-cause mortalityGeneral urban population (>300,000); no specific vulnerability subgroup
(Boudreault et al. 2023)1981–2019CMA-level (Montreal)Extreme heat (percentile-based)Daily mortality deviation (excess deaths above seasonal baseline)General urban population, Montreal CMA; no age stratification reported
(Boudreault et al. 2024)1998–2019Metropolitan area (Montreal and Quebec CMA)Extreme heat exposure (continuous Temp metrics, lag-based)All-cause daily mortality ratesGeneral urban population; elderly (≥ 65 years) highlighted as key subgroup
(Côté et al. 2024)2001–2018Province-wide (Quebec), stratified by CDD climate regionsHeatwave: 3-day moving average of daily Tmax exceeding percentile thresholdDaily all-cause mortality rate per 10,000General population (Quebec); stratified by CDD regions; elderly subgroup
(Jian et al. 2023)2006–2015Metropolitan area, Perth (21 SA3 units)Heatwave: EHF > 0 (80th percentile cut-off)EHF threshold exceedance; ED attendance for heat-sensitive conditionsGeneral urban and child populations; children under 5 years as a vulnerable subgroup
(Ke et al. 2023)1991–202011 climate regions (1 × 1 km grid, prefecture-level)Extreme heat via WBGT and accumulated heat stress thresholdsHeat-related ambulance calls; diagnosed heat strokePrefecture-level populations; outdoor workers; elderly adults
(Kim and Kim 2022)2016–2018City-level (Daegu, eight districts)Heatwave (daily temp indicators, HI-based)Heat-related deaths: hypertensive diseases, IHD, and CEVUrban metropolitan population; elderly, young children (< 5 years), socially isolated elderly
(Nishimura et al. 2021)2014–2019City-level (Nagoya, 16 wards)Heatwave: daily working-time Tavg; WBGT threshold exceedanceAmbulance transport for heat-related illnessGeneral population exposed to heat; ward-level variation examined
(Ohashi et al. 2023)2009–2019City-level (1 km grid, Tokyo 23 wards)Heatwave/heat exposure (daily Tmax, lagged cumulative indices)Mortality due to IHD and CEVUrban population (Tokyo); elderly adults (≥ 65 years) as primary risk group
(Schachtschneider et al. 2024)2015–2021National level (1° × 1 ° grid)Heatwave: monthly maximum 2 m air tempAll-cause mortality per 100,000 populationEntire national population; future climate scenario projections included
(Sophia et al. 2025)2004–2015District-level (Pune)Heavy rainfall: wet weeks (0.5–150 mm/week); rainfall flush events (>150 mm/week)Dengue mortalityUrban areas (~80% from Pune city); general population; vector-borne disease context
(Toure et al. 2025)2017–2022Regional (Matam)Heatwave: ≥ 3 consecutive days above 90th percentileHospital admissions for heat-sensitive pathologiesGeneral population exposed to heatwaves; single regional hospital catchment
(Wang et al. 2024)2011–2020National level (district-level resolution)Heatwave: warm days > 20 °CHeat-related mortalityUrban and rural populations; district-level heterogeneity examined
(Wang et al. 2023)2016–2018Nationwide (SA1 resolution, > 57,000 units)Heatwave: EHF (3-day Tavg vs. 95th percentile Tmax)Prevalence of heat-vulnerable chronic diseasesGeneral urban and rural populations; persons with pre-existing heat-sensitive conditions
(Wang et al. 2019)2012–2014City-level (500 m resolution, seven cities)Heatwave: days with Tmax > 35 °CDaily heatstroke casesUrban populations in southern China, high ambient temp settings
(Wang et al. 2025)2014–2016City-level (1 km grid, Wuhan)Extreme heat > 95th percentile (30.57 °C); extreme cold < 5th percentile (3.41 °C)PTB (< 37 weeks gestation)311,972 pregnant women; dual thermal exposure framework
(Wertis et al. 2023)2016–2019City-level (six US cities)Extreme heat: daily temp indicators (continuous metrics, no formal heatwave definition)Mental and behavioral disorders (ER visits)Urban youth aged 5–24 years; socio-demographic vulnerability highlighted
(Xu et al. 2024)2014–2019City-level (southern China)Heatwave: Tmax ≥ 35 °CHeatstroke incidenceUrban population; no specific subgroup stratification reported

Study and clinical characteristics of included studies.

Studies ordered alphabetically by first author surname. Full methodological details, validation strategies, and feature importance analyzes are provided in Supplementary files S2, S3. CEV, Cerebrovascular Disease; CDD, Cooling Degree Days; CMA, Census Metropolitan Area; ED, Emergency Department; EHF, Excess Heat Factor; ER, Emergency Room; EWE, Extreme Weather Event; IHD, Ischemic Heart Disease; PTB, Preterm Birth; SA1, Statistical Area Level 1; SA3, Statistical Area Level 3; Temp, Temperature; Tavg, Average Temperature; Tmax, Maximum Temperature; Tmean, Mean Temperature; WBGT, Wet Bulb Globe Temperature.

According to World Bank income classifications (World Bank, 2024), most studies (n = 13) were conducted in HICs, with three studies from upper-middle-income countries (UMICs), specifically China, and two studies from LMICs, India and Senegal. Data collection occurred at multiple geographic scales, including city or metropolitan, district, regional and national levels. City or metropolitan-level analyzes were most common (Wang et al., 2019; Wertis et al., 2023; Boudreault et al., 2024; Ohashi et al., 2023; Xu et al., 2024; Kim and Kim, 2022; Boudreault et al., 2023; Nishimura et al., 2021; Jian et al., 2023; ; Wang et al., 2025), followed by national (Schachtschneider et al., 2024; Wang et al., 2023), district (Sophia et al., 2025; Wang et al., 2024), and regional (Ke et al., 2023; Toure et al., 2025) scales.

3.2 Extreme weather event characterization

The analysis of EWE definitions across the 18 included studies revealed a predominant focus on heat-related exposures, with 17 studies (94.4%) examining various forms of thermal stress and only one study investigating extreme rainfall, as detailed in Table 3. Absolute temperature thresholds were applied (n = 2), with high-temperature or heatwave days defined as daily maximum temperatures exceeding 35 °C in urban Chinese settings (Wang et al., 2019; Xu et al., 2024). Percentile-based approaches were used (n = 4), with extreme heat defined as exceedance of the 85th, 90th, 95th, or 98th percentiles of local temperature distributions (; Ke et al., 2023; Toure et al., 2025; Wang et al., 2025). Duration-based criteria were used (n = 4) to distinguish single-day extreme heat events from multi-day heatwaves, with heatwave conditions defined by either two () or 3 consecutive days (Toure et al., 2025; Wang et al., 2025; Xu et al., 2024) above the specified threshold.

In contrast, continuous daily temperature metrics or warm-day counts were modeled (n = 3), without specifying a formal heatwave definition (Wertis et al., 2023; Kim and Kim, 2022; Boudreault et al., 2024). Heat exposure was assessed at a daily temporal resolution (n = 17), with studies evaluating concurrent effects or short lag periods of up to 3 days (n = 5) (Wang et al., 2019; Xu et al., 2024; Ohashi et al., 2023; Wang et al., 2024, 2025), and studies incorporating extended lag structures of up to 7 days (n = 3) (Ke et al., 2023; Toure et al., 2025; Boudreault et al., 2024). Longer or non-standard temporal dependencies beyond conventional lag structures were reported in two studies (Nishimura et al., 2021; Sophia et al., 2025). Regional patterns were evident, with studies conducted in temperate settings, including Canada and Germany (Wang et al., 2024; Boudreault et al., 2024; Côté et al., 2024) more frequently employing percentile-based or moderate absolute thresholds, whereas studies conducted in warmer climatic settings, including parts of China, Japan, and Australia (Wang et al., 2019; Xu et al., 2024; Ke et al., 2023; Wang et al., 2023), commonly applied higher absolute cut-points aligned with established heat-health thresholds. The single rainfall-focused study (Sophia et al., 2025), defined exposure using weekly cumulative precipitation thresholds, identifying wet weeks (0.5–150 mm/week) and rainfall flush events (>150 mm/week) with a 2-month lead forecasting framework, highlighting the distinct methodological requirements for non-thermal EWEs. One study uniquely examined dual thermal exposure (Wang et al., 2025), defining extreme heat above the 95th percentile (30.57 °C) and extreme cold below the 5th percentile (3.41 °C), providing a framework for comprehensive temperature-health modeling across both ends of the thermal spectrum.

3.3 Health outcomes and target population

Table 3 shows that the health outcomes evaluated across the included studies were primarily acute mortality and morbidity associated with EWEs, but also included longer-term or non-acute outcomes, such as chronic disease prevalence and adverse birth outcomes. All-cause mortality was the most frequently examined outcome (n = 6), and was assessed at daily or monthly resolutions using national or regional death registries (Boudreault et al., 2024; Schachtschneider et al., 2024; Boudreault et al., 2023; Côté et al., 2024; ; Wang et al., 2024). Cause-specific mortality was examined in two studies focusing on cardiovascular and cerebrovascular deaths defined using International Classification of Diseases (ICD) coded records (Ohashi et al., 2023; Kim and Kim, 2022). Acute heat-related morbidity outcomes were assessed in six studies, including heatstroke incidence reported by surveillance systems (Wang et al., 2019; Xu et al., 2024), heat-related ambulance calls (Ke et al., 2023; Nishimura et al., 2021), and emergency department visits for heat-sensitive conditions (Jian et al., 2023; Toure et al., 2025). Further outcome domains were examined (n = 4), including mental and behavioural disorder-related emergency visits (Wertis et al., 2023), PTB among pregnant women (Wang et al., 2025), dengue mortality during monsoon periods (Sophia et al., 2025), and the prevalence of heat-vulnerable chronic diseases (Wang et al., 2023). Studies relied on administrative health data sources, including death registries, hospital records, and emergency services databases, with outcome temporal resolution ranging from daily (n = 16), with one study each reporting monthly and annual incidence measures.

Population characteristics and vulnerability patterns vary significantly across studies, reflecting differences in geographic contexts and research objectives. Age-stratified analyzes were observed (n = 7), with particular emphasis on older adults aged 65 years and above (Ke et al., 2023; Boudreault et al., 2024; Ohashi et al., 2023; Côté et al., 2024) as well as pediatric and youth groups, including children under 5 years (Kim and Kim, 2022; Jian et al., 2023) and adolescents aged 5-24 years (Wertis et al., 2023). Studies also focused on specific vulnerable subpopulations (n = 4), including pregnant women (Wang et al., 2025), outdoor workers (Ke et al., 2023), socially isolated elderly individuals (Kim and Kim, 2022), and populations with pre-existing heat-sensitive conditions (Wang et al., 2023).

3.4 Environmental drivers and predictive variables

Across the included studies, environmental predictors comprised meteorological variables, air quality indicators, and derived exposure metrics, with substantial variation in their definition, operationalization, and incorporation into predictive models, as summarized in Table 4, and further detailed in Supplementary files S2, S3.

Table 4

ReferencesPrimary predictorsSES/demographic included?ML modelsBest model and performanceValidationROB (PROBAST)
()Tmean, relative temp intensity, percentile-based heat metricsNoClassification tree, conditional tree, bagging, RF, and boosting (20 variants)Boosting/RF (ROSE); sensitivity ≈ 94% for rare high-mortality HWsMonte Carlo CV (100 runs; 2/3–1/3 split); class-imbalance correctionP: Low
Pr: Low
O: Low
A: Unclear
Overall: unclear
(Boudreault et al. 2023)Tmax, humidity, wind speed, air pollution (PM2.5, O3); DLNM lag structuresNoDT, RF, GBM, SLP, MLP, LSTM (n = 6)GBM; best RMSE, MAE, R2 vs. statistical baselinesTemporal 70/30 split; five-fold CV tuning; HW-based partial validationP: Low
Pr: Low
O: Low
A: Low
Overall: low
(Boudreault et al. 2024)Tmax, humidity, precipitation, pressure, wind, air pollution; lagged & aggregatedNoGAM, RF, GBM, MLP, and LSTM (n = 5)GAM; R2=6–7%; matched or outperformed ML modelsTemporal 70/30 hold-out; CV-based hyperparameter tuningP: Low
Pr: Low
O: Low
A: Low
Overall: low
(Côté et al. 2024)Tmax (3-day moving average), CDD region, vegetation (NDVI), and air qualityYes (mixed: SES + infrastructure)AutoGluon, GP, deep GPGP; best peak HW mortality capture; uncertainty-aware predictionsLOYO temporal validation; multiple test sets; RMSE stratified by heat intensityP: Low
Pr: Low
O: Low
A: Unclear
Overall: unclear
(Jian et al. 2023)EHF, Tmax, air pollution; lagged exposure metrics; NDVI (vegetation)Yes (mixed: SES + demographic)DT, RF, and GRFRF; R2=0.953 for ED HW attendance prediction70/30 split; 500-fold CV; temporal construct validationP: Low
Pr: Low
O: Low
A: Unclear
Overall: unclear
(Ke et al. 2023)WBGT, accumulated heat stress, Tmax, humidity, and wind speedYes (mixed: demographic + infrastructure)XGB (four variants: national/regional; with/without HW features)Regional XGB with HW features; adj. R2=0.98Repeated 10-fold CV; train/test evaluationP: Low
Pr: Low
O: Low
A: Low
Overall: low
(Kim and Kim 2022)Tmax, HI, humidity; grid-search optimized featuresYes (mixed: demographic + SES)RF (grid-search optimized)RF; AUC = 0.86, ACC = 90.3%, F1 = 0.95; SHAP-based interpretabilityHold-out validation; GridSearchCV tuningP: Low
Pr: Low
O: Low
A: Low
Overall: low
(Nishimura et al. 2021)Tavg (daily working-time), WBGT; ward-level environmental metricsYes (demographic)LSTM, RFR (n = 2)LSTM; R2≈0.80; lowest residual error vs. RFRLOYO temporal CV; ward-level evaluationP: Low
Pr: Low
O: Low
A: Unclear
Overall: unclear
(Ohashi et al. 2023)Tmax, humidity, precipitation, pressure; lagged and cumulative indicesNoLightGBM (n = 1)LightGBM; RMSE = 0.369, MAE = 0.290 for IHD mortality10-fold CV (90/10 split)P: Low
Pr: Low
O: Low
A: Low
Overall: low
(Schachtschneider et al. 2024)Monthly maximum 2 m air temp (reanalysis-derived); no composite indicesNoESN (n = 1)ESN; RMS = 1.7; accurate under future climate scenariosTemporal hold-out; ensemble validation (25 ESNs)P: Low
Pr: Low
O: Low
A: Low
Overall: low
(Sophia et al. 2025)Weekly cumulative precipitation (wet weeks; flush events); Tmean, humidityNoRFR (n = 1)RFR; r = 0.77, NRMSE = 0.52; 2-month dengue early warningTemporal train/test split; K-fold CVP: Low
Pr: Low
O: Low
A: Low
Overall: low
(Toure et al. 2025)Tmax, humidity; quantile and cumulative heat indicesYes (demographic)RF, XGBRF; R2=0.51–0.72; captured delayed HW hospitalization effects80/20 split; 10-fold CV; grid search; 1,000-iteration bootstrappingP: Low
Pr: Low
O: Low
A: Low
Overall: low
(Wang et al. 2024)Tmax, Tmean (daily and weekly); no composite indices reportedNoNNs (linear and exponential architectures; ensemble-based; n = 20)NN ensemble; R2=0.83; district-level heat mortality and lag effectsTemporal 80/20 split; ensemble validation (20 models)P: Low
Pr: Low
O: Low
A: High Overall: high
(Wang et al. 2023)EHF, Tmax, air quality (NDVI, satellite-derived); seasonal summariesYes (mixed: SES + built environment)RFR (n = 1)RFR; R2 ≤ 0.99 (urban), ≥ 0.75 (rural); outperformed OLS and MGWR80/20 train–test split; grid-search tuning; robustness checks (OLS, MGWR)P: Low
Pr: Low
O: Low
A: Low
Overall: low
(Wang et al. 2019)Tmax, humidity, wind speed, NDVI; lagged temperature featuresYes (mixed: SES + demographic)RF (n = 1)RF; R2 = 0.70; outperformed LR; captured non-linear heatstroke risk90/10 split; 10-fold CV; sensitivity analysis; Bland–Altman calibrationP: Low
Pr: Low
O: Low
A: Unclear
Overall: unclear
(Wang et al. 2025)Tmax, Tmin, humidity, air pollution (PM2.5); extreme day counts (both thermal ends)Yes (mixed: demographic + SES)XGB (n = 1)XGB; AUC=0.93; non-linear temperature thresholds for PTB risk zones5-fold CV; Bayesian optimizationP: Low
Pr: Low
O: Low
A: Unclear
Overall: unclear
(Wertis et al. 2023)Tmax, humidity, NDVI; DLNM sensitivity; daily temp indicatorsYes (mixed: demographic + SES)GLM, GAM, RF, XGB (n = 4)GAM; RMSE = 4.96, MAE = 3.59; lowest prediction error80/20 split; five-fold CV; DLNM sensitivity analysisP: Low
Pr: Low
O: Low
A: Low
Overall: low
(Xu et al. 2024)Tmax, humidity, HI; heatwave indicatorsNoRegression DT, RF, GBDT, linear SVR, LSTM, and ARIMA (n = 6)RF; RMSE = 3.57, R2 = 0.82; lowest error; temporal out-of-year calibrationTemporal train–test split (2014–2018/2019); grid-search tuningP: Low
Pr: Low
O: Low
A: Low
Overall: low

Synthesis of predictor characteristics, modeling frameworks, performance metrics, validation strategies, and risk of bias across the included studies.

ROB assessed using PROBAST domains: P, Participants; Pr, Predictors; O, Outcome; A, Analysis. Overall ROB reflects the highest-concern domain; full domain-level detail is provided in Supplementary files S2, S3.

ACC, Accuracy; ARIMA, Autoregressive Integrated Moving Average; AUC, Area Under the Curve; CV, Cross-Validation; Dem., Demographic; DLNM, Distributed Lag Non-linear Model; DT, Decision Tree; ESN, Echo State Network; F1, F1-Score; GAM, Generalized Additive Model; GBM, Gradient Boosting Machine; GBDT, Gradient Boosted Decision Tree; GLM, Generalized Linear Model; GP, Gaussian Process; GRF, Generalized Random Forest; HI, Heat Index; HW, Heatwave; LOYO, Leave-One-Year-Out; LR, Logistic Regression; LSTM, Long Short-Term Memory; MAE, Mean Absolute Error; MGWR, Multiscale Geographically Weighted Regression; MLP, Multilayer Perceptron; MLR, Multiple Linear Regression; NR, Not Reported; NN: Neural Network; NRMSE, Normalized RMSE; OLS, Ordinary Least Squares; RF, Random Forest; RFR, Random Forest Regression; RMSE, Root Mean Square Error; ROB, Risk of Bias; ROSE, Random Over-Sampling Examples; SES, Socioeconomic Status; SHAP, Shapley Additive Explanations; SLP, Single-Layer Perceptron; SVR, Support Vector Regression; Tmax, Maximum Temperature; XGB, Extreme Gradient Boosting. Bold values denote the overall risk-of-bias (ROB) assessment for each study.

3.4.1 Environmental predictor categories and usage frequency

Across all 18 included studies, temperature variables were consistently used as environmental predictors to model health outcomes associated with EWEs, as shown in Figure 5A. Temperature information was incorporated directly as daily, weekly, or monthly metrics or indirectly through composite heat stress and heatwave indices. The representation of temperature varied across studies, as shown in Figure 5B, but these counts reflect variable-level frequency and are not directly comparable to the study-level classification. Seven studies relied exclusively on raw meteorological measurements, including daily maximum, minimum, and mean temperature values (Wang et al., 2019; Ohashi et al., 2023; Wang et al., 2024; Boudreault et al., 2023, 2024; Sophia et al., 2025; Wang et al., 2025). Two studies used only derived or non-raw temperature representations, such as relative intensity metrics or reanalysis-based fields (Schachtschneider et al., 2024; ). The remaining nine studies incorporated both raw and derived or composite indices, including the Excess Heat Factor (EHF), Heat Index (HI), Wet Bulb Globe Temperature (WBGT), accumulated heat load, or percentile-based extreme thresholds (Ke et al., 2023; Wertis et al., 2023; Xu et al., 2024; Kim and Kim, 2022; Nishimura et al., 2021; Jian et al., 2023; Côté et al., 2024; Toure et al., 2025; Wang et al., 2023), as shown in Figure 5C. The temporal resolution of environmental predictors varied across studies, with daily data used in 15 studies (83%), followed by weekly (n = 2, 11%) and monthly (n = 1, 6%) data, as shown in Figure 5D.

Figure 5

Relative humidity (RH) was the second most commonly used meteorological variable, appearing in 11 studies (61%), and was included either directly as RH measurements or indirectly through composite heat stress indices such as HI, WBGT, humidex, and dew point temperature (Wang et al., 2019; Ke et al., 2023; Wertis et al., 2023; Xu et al., 2024; Kim and Kim, 2022; Boudreault et al., 2024; Ohashi et al., 2023; Sophia et al., 2025; Boudreault et al., 2023; Wang et al., 2025; Toure et al., 2025). In six studies, humidity was incorporated using variable-specific temporal structures, including short-term lags (0–5 days) and longer lag periods identified through cross-correlation or distributed lag models (Wang et al., 2019; Boudreault et al., 2024; Ohashi et al., 2023; Xu et al., 2024; Boudreault et al., 2023; Sophia et al., 2025). Humidity-related variables ranked among the top three predictors in seven studies (Wang et al., 2019; Ke et al., 2023; Wertis et al., 2023; Boudreault et al., 2024; Xu et al., 2024; Kim and Kim, 2022; Sophia et al., 2025). Other meteorological variables were used less frequently, including wind speed (n = 4) (Wang et al., 2019; Ke et al., 2023; Boudreault et al., 2024, 2023), precipitation (n = 3) (Boudreault et al., 2024; Ohashi et al., 2023; Sophia et al., 2025), and atmospheric pressure (n = 2) (Boudreault et al., 2024; Ohashi et al., 2023), as shown in Figure 5A.

Air quality variables encompassing particulate matter and gaseous pollutants were integrated in six studies (33%), with these studies incorporating both meteorological and air quality predictors in the same models (Boudreault et al., 2023; Côté et al., 2024; Wang et al., 2023; Jian et al., 2023; Wang et al., 2025; Boudreault et al., 2024). The most commonly included air quality variables were PM2.5 (n = 5) (Boudreault et al., 2023, 2024; Côté et al., 2024; Jian et al., 2023; Wang et al., 2025) and O3 (n = 5) (Boudreault et al., 2023, 2024; Wang et al., 2023; Jian et al., 2023; Wang et al., 2025), followed by NO2 and SO2, each reported in four studies (Boudreault et al., 2023, 2024; Wang et al., 2023; Jian et al., 2023), as summarized in Table 4. Vegetation and land cover indicators, most commonly the Normalized Difference Vegetation Index (NDVI), were employed in three studies (17%) (Wang et al., 2019; Wertis et al., 2023; Côté et al., 2024), as shown in Figure 5A.

3.4.2 Environmental predictor mapping

Figure 6 maps the dominant environmental predictor categories against modeled health outcomes across the 18 included studies. Temperature-based variables were used across five of the seven outcome categories and were primarily associated with mortality and acute morbidity outcomes. Daily Tmax predicted all-cause and heat-related mortality in one study (Wang et al., 2024), cardiovascular and cerebrovascular mortality in another (Ohashi et al., 2023), and heatstroke incidence in two studies (Xu et al., 2024; Wang et al., 2019). Monthly Tmax was used in one study to model all-cause mortality (Schachtschneider et al., 2024). Composite heat stress indices, including WBGT, EHF, and HI, were used as predictors of ambulance dispatch in three studies (Nishimura et al., 2021; Toure et al., 2025; Ke et al., 2023) and emergency department attendance in two (Wertis et al., 2023; Jian et al., 2023), with another study employing composite heat stress indices for heatstroke prediction (Xu et al., 2024).

Figure 6

Temperature combined with air quality indicators characterized mortality prediction in two Canadian studies (Boudreault et al., 2023, 2024), while two additional studies incorporated air pollution alongside other predictor types across emergency department attendance and PTB outcomes (Jian et al., 2023; Wang et al., 2025). Temperature combined with socioeconomic and demographic variables was used across five of seven outcome categories, as shown in Figure 6, and was reported in eight studies (Côté et al., 2024; Wang et al., 2023; Toure et al., 2025; Kim and Kim, 2022; Wang et al., 2019; Wertis et al., 2023; Jian et al., 2023; Wang et al., 2025). Beyond heat-related exposure, the evidence base is substantially narrowed. Weekly cumulative rainfall served as the primary environmental predictor in a single vector-borne disease study examining dengue mortality in India, with humidity included as a covariate, and rainfall was further categorized into wet weeks (0.5-150 mm/week) and high-intensity rainfall flushes (>150 mm/week) (Sophia et al., 2025). However, no included study explicitly modeled any health outcomes using flood-related environmental exposure.

3.4.3 Socioeconomic and demographic variables

Beyond environmental predictors, 11 of the 18 studies (61%) incorporated at least one non-environmental predictor, including demographic, socioeconomic, and temporal variables. Of these, 10 studies explicitly included demographic and/or socioeconomic indicators, while one study incorporated temporal predictors without socioeconomic or demographic variables (Xu et al., 2024). Demographic variables primarily capture population structure, including sex, proportions of older adult populations, children, and age-stratified population counts, and are commonly used to stratify heat-related health outcome (Ke et al., 2023; Nishimura et al., 2021; Toure et al., 2025; Jian et al., 2023; Kim and Kim, 2022). Socioeconomic indicators derived from census or administrative sources include income, education, employment status, and area-level composite indices (Côté et al., 2024; Wang et al., 2019; Jian et al., 2023; Wang et al., 2023). Infrastructure-related indicators, such as air conditioning availability, healthcare resources, and cooling facilities, were incorporated into two studies as proxies for adaptive capacity (Ke et al., 2023; Côté et al., 2024).

However, in three studies, demographic or socioeconomic variables were identified as the most influential predictors in feature importance analysis, based on SHAP values or permutation-based importance metrics. These include young and older adult population counts in heat mortality prediction (Kim and Kim, 2022), area-level socioeconomic index in the stratified young children model of emergency visits in Australia (Jian et al., 2023), and age composition and sex ratio in emergency visits for mental and behavioural disorders in the United States (Wertis et al., 2023). In the latter study, SHAP analyzes indicated that demographic predictors had larger magnitude contributions to the model output than temperature and humidity; the authors interpreted this as sociodemographic factors contributing alongside environmental parameters rather than acting independently. In two further studies, socioeconomic and demographic variables ranked above environmental predictors only in specific subgroups, including urban populations and higher heat intensity strata, but not consistently across all model configurations (Wang et al., 2023; Côté et al., 2024). Across all five studies (Kim and Kim, 2022; Jian et al., 2023; Wertis et al., 2023; Wang et al., 2023; Côté et al., 2024), the population age structure and socioeconomic position were among the most influential predictors in at least one model configuration, particularly for heat-sensitive subgroups.

3.5 Machine learning approaches

The distribution of ML algorithms evaluated across the 18 included studies shows patterns in both evaluation frequency and performance outcomes, as detailed in Table 4.

3.5.1 Algorithm selection and baseline comparison

RF was the most frequently evaluated algorithm, appearing in 12 studies (67%), and was identified as the best-performing model in seven studies (58%) (Wang et al., 2019; Xu et al., 2024; Wang et al., 2023; Jian et al., 2023; ; Toure et al., 2025; Kim and Kim, 2022), based on reported performance metrics such as R2, RMSE, and AUC. As illustrated in Figure 7, RF shows a higher proportion of best-model selections relative to its evaluation frequency. Boosting-based variants, including XGBoost, GBM, and LightGBM, were evaluated in nine studies (Ke et al., 2023; Wertis et al., 2023; Xu et al., 2024; Boudreault et al., 2023, 2024; Wang et al., 2025; Toure et al., 2025; Côté et al., 2024; ), and were identified as optimal in three of those studies (33%) (Ke et al., 2023; Boudreault et al., 2023; ), across mortality, morbidity, and health service utilization outcomes. Notably, a single rainfall-focused study (Sophia et al., 2025) employed RFR to predict dengue incidence in relation to extreme precipitation events, achieving moderate predictive performance (r = 0.77, NRMSE = 0.52) with a 2-month forecast lead time. In contrast to heat-focused studies, which predominantly used daily temporal resolution and short lag structures, this study was conducted at weekly resolution with longer forecast lead times.

Figure 7

Furthermore, DL approaches, including LSTM (n = 4) (Boudreault et al., 2024, 2023; Xu et al., 2024; Nishimura et al., 2021), NNs (n = 3) (Boudreault et al., 2024; Nishimura et al., 2021; Wang et al., 2024), and ESN (n = 1) (Schachtschneider et al., 2024), were evaluated less frequently and were rarely identified as best-performing models. LSTM-based models showed advantages primarily in temporal forecasting tasks, such as heat-related illness incidence prediction, where they outperformed RF under LOYO validation (Nishimura et al., 2021). However, in multi-model benchmarking studies, DL approaches generally do not outperform ensemble tree-based methods (Boudreault et al., 2023). Baseline model inclusion varied across studies, with 11 of the 18 studies (61%) (Wang et al., 2019; Wertis et al., 2023; Boudreault et al., 2024; Xu et al., 2024; Boudreault et al., 2023; Wang et al., 2023; Nishimura et al., 2021; Jian et al., 2023; ; Toure et al., 2025; Wang et al., 2024) incorporating traditional statistical comparators. GAM were the most common baseline (n = 4) (Wertis et al., 2023; Boudreault et al., 2024, 2023; Toure et al., 2025), followed by linear regression variants (n = 4) (Wang et al., 2019, 2023; Jian et al., 2023; Wang et al., 2024), and DLNM (n = 2) (Xu et al., 2024; Boudreault et al., 2023). GAM demonstrated competitive or superior performance in two studies (Wertis et al., 2023; Boudreault et al., 2024), while ensemble tree-based methods outperformed statistical comparators in studies using multivariable comparative frameworks (Wang et al., 2023; Jian et al., 2023). GAM also achieved out-of-sample explanatory power comparable to ML approaches in mortality-focused modeling (Boudreault et al., 2023), as detailed in Table 4.

3.5.2 Feature engineering and model configuration

Figure 5C summarizes feature engineering approaches across all 18 included studies, which formed a central component of model development. Raw environmental variables were transformed into derived predictors in most studies (n = 16, 89%), including representations of delayed, cumulative, and extreme exposure. Approaches include the creation of lagged temperature and humidity features over varying temporal windows (Wang et al., 2019; Boudreault et al., 2024; Ohashi et al., 2023), rolling or weighted temperature averages (Ke et al., 2023; Ohashi et al., 2023), and percentile-based extreme heat indicators or implicitly learned high-temperature response functions (Xu et al., 2024; Wang et al., 2024). Additional engineered representations of environmental variables include temperature variability metrics and day-to-day temperature changes (Boudreault et al., 2024; Wertis et al., 2023), composite thermal indices including HI, WBGT, and EHF used as model input features (Ke et al., 2023; Schachtschneider et al., 2024), and spatially aggregated or gridded temperature features (Wang et al., 2024). Explicit feature selection or feature importance analysis was conducted in 83% of the studies (n = 15). Methods included Boruta-based selection (Wang et al., 2019; Xu et al., 2024; Ohashi et al., 2023), node purity-based and permutation importance measures (Wang et al., 2019; Boudreault et al., 2024, 2023; Jian et al., 2023; Wang et al., 2023; Sophia et al., 2025; ), and SHAP-based explainability frameworks (Ke et al., 2023; Wertis et al., 2023; Ohashi et al., 2023; Xu et al., 2024; Kim and Kim, 2022; Côté et al., 2024; Wang et al., 2025; Toure et al., 2025).

Hyperparameter optimization was reported in 72% of the studies (n = 13), with grid search being the most frequently applied method, followed by Bayesian optimization and randomized search. Ensemble strategies are pervasive, with RF relying on bootstrap aggregation and boosting-based models that employ sequential error correction. Across studies, feature importance analyzes have identified temperature-related variables as among the most influential predictors of health outcomes, with maximum, mean, and heat index-based temperature metrics frequently ranking among the most important features (Wang et al., 2019; Xu et al., 2024; Kim and Kim, 2022). Sociodemographic variables ranked above environmental predictors in three studies and in specific subgroups in two further studies (see Section 3.4.3). Lagged temperature effects were identified as influential predictors, with 1–3 day lags frequently ranked among the most important features (Boudreault et al., 2024; Ohashi et al., 2023). A detailed description of the feature selection and importance analysis is provided in Supplementary files S2, S3.

3.5.3 Validation frameworks and performance assessment

The validation strategies varied across the 18 included studies, with some employing more than one validation approach. K-fold cross-validation (CV) was the most frequently applied internal validation approach, used as a standalone method in five studies (28%), employing the 5-fold or 10-fold schemes detailed in Table 4. Temporal validation was applied in eight studies (44%) using chronological hold-out splits or LOYO approaches. More intensive resampling strategies, including Monte Carlo CV and repeated validation, were employed in two studies (11%) (; Jian et al., 2023). Four studies (22%) relied solely on random train-test splits without CV or temporal structuring (Kim and Kim, 2022; Wang et al., 2023; Boudreault et al., 2023; Wang et al., 2025). External validation was not reported in any of the included studies; all studies relied on internal validation strategies without evaluating model performance on independent datasets across different regions or time periods.

Performance metrics varied by outcome type, with distinct patterns observed between the classification and regression studies. Classification studies mostly used AUC-ROC (n = 6 studies) and accuracy measures (n = 5 studies), with reported AUC values ranging from 0.80 to 0.93 (Kim and Kim, 2022; Wang et al., 2025). Regression-based studies emphasized R2 (n = 10 studies), RMSE (n = 8 studies), and MAE (n = 6 studies), with reported values ranging from modest population-level mortality prediction (R2≈0.06 − 0.08) (Boudreault et al., 2024) to high-precision morbidity modeling (R2>0.90) (Jian et al., 2023; Xu et al., 2024). One study additionally reported a Bland-Altman agreement analysis alongside regression metrics (Wang et al., 2019).

3.5.4 Regional patterns in model performance

The best-reported predictive performance stratified by country and income classification across the nine countries represented in this review is illustrated in Figure 8. Where multiple studies were conducted within the same country, the highest reported performance values are presented for consistency; full study-level performance details are provided in Supplementary files S2, S3. Performance values are not directly comparable across countries, as studies differed in outcome type and evaluation metrics. Among HICs, Japan reported an adjusted R2 of 0.98 using XGBoost (Ke et al., 2023), whereas in Australia, R2 = 0.95 was achieved using RF (Jian et al., 2023). In the US, boosting-based models have reported a sensitivity of 0.94 (). In South Korea, RF achieved an AUC of 0.86 (Kim and Kim, 2022), and in Germany, NNs reported an R2 of 0.83 (Wang et al., 2024). In contrast, in Canada, GAM reported an R2 of 0.07 (Boudreault et al., 2024).

Figure 8

Among the UMICs and LMICs, RF achieved an R2 of 0.82 in China (Wang et al., 2023), whereas in India, RFR reported a Pearson correlation of r = 0.77 (Sophia et al., 2025), and in Senegal, RF achieved an R2 of 0.72 (Toure et al., 2025). Overall, RF was the most frequently identified best-performing algorithm, appearing in six of the nine representative countries.

4 Discussion

This review synthesized evidence from 18 studies across nine countries examining ML approaches for predicting health outcomes associated with EWEs. The evidence base was predominantly derived from high-income settings (72%, n = 13) and focused almost exclusively on heat-related exposure (94%, n = 17). The following subsections interpret these findings in relation to model transferability, thematic gaps, and priorities for future research.

4.1 Predictor-outcome and context sensitivity

The predictor and health outcome patterns summarized in Figure 6 reveal systematic alignments between exposure type and outcome specificity. The dominance of composite heat stress indices, such as WBGT and EHF, for acute morbidity outcomes, including heatstroke and ambulance calls, is consistent with their ability to capture the combined thermal burden of temperature, humidity, and radiant heat. Therefore, these indices more closely approximate the physiological pathways of thermal injury than raw temperature alone, as sustained thermal load impairs thermoregulatory capacity by reducing evaporative cooling efficiency, increasing physiological heat strain and elevating the risk of heat-related morbidity and mortality, particularly among populations with limited adaptive capacity (World Health Organization, 2024; Ke et al., 2023; Nishimura et al., 2021). Temperature was used to characterize both acute exposure during extreme heat events and cumulative thermal stress, reflecting its central role across the diverse heat-related health outcomes examined in the included studies.

The inclusion of humidity further reflects its role in heat stress physiology, particularly in humid environments where thermal strain accumulates more rapidly. Vegetation indices were employed in three studies to represent the environmental buffering capacity and urban heat island effects (Wang et al., 2019; Wertis et al., 2023; Côté et al., 2024). In contrast, percentile-based temperature thresholds and daily Tmax metrics dominate population-level mortality models, reflecting the threshold-dependent structure of heat-health relationships at an aggregate scale (Boudreault et al., 2023; Côté et al., 2024). The identification of socioeconomic or demographic variables as the highest-ranked predictors in three studies further indicates that, in heterogeneous urban populations, the vulnerability context materially modifies environmental predictors and health outcome relationships and should be considered integral rather than optional within predictive frameworks (Kim and Kim, 2022; Jian et al., 2023; Wertis et al., 2023). This suggests that environmental exposure alone may not adequately capture risk without accounting for population-level susceptibility and adaptive capacity. While the reviewed studies do not model these exposure-response pathways mechanistically or estimate causal environmental-health relationships directly, the alignment between predictor selection and established physiological and epidemiological mechanisms provides biological plausibility for the observed associations and underscores the value of these predictors as candidates for integration into future mechanistic and hybrid modeling frameworks.

The observed transition from temperature-based predictors for heat-related outcomes to precipitation-based predictors for vector-borne diseases, together with the complete absence of flood-specific predictor evidence, underscores that predictor selection must be outcome-specific and context-sensitive, even where rainfall thresholds approach flood-adjacent conditions, as in the single rainfall-based study (Sophia et al., 2025). Expanding the evidence base beyond heat will require coordinated investment in environmental monitoring infrastructure and geocoded health surveillance systems capable of linking exposure and outcomes at appropriate spatial and temporal resolutions.

4.2 Model performance, generalization, and geographics disparities

Ensemble tree-based methods, particularly RF, demonstrated the strongest overall performance across the evidence base and were identified as the best-performing models in seven studies (Wang et al., 2019; Xu et al., 2024; Wang et al., 2023; Jian et al., 2023; ; Toure et al., 2025; Kim and Kim, 2022). However, the algorithm performance was not uniform, as in two studies, GAM achieved superior or joint-best performance relative to ML models, indicating that conventional regression approaches can retain competitive utility when predictor-outcome relationships are relatively stable or data complexity is limited (Wertis et al., 2023; Boudreault et al., 2024). In other studies, GB variants, including LightGBM and XGBoost, emerged as the strongest performers (Ohashi et al., 2023; Ke et al., 2023). This pattern of conditional performance suggests that algorithm selection should be guided by data structure, outcome type, and modeling objective rather than a default preference for any single approach. This robustness of RF, where it performed well, likely reflects three structural advantages: the ability to capture non-linear exposure-response relationships, resilience to multicollinearity among environmental predictors, and the generation of interpretable feature importance measures (Ciecierski-Holmes et al., 2022). The limited performance advantage of DL approaches in this review contrasts with their dominance in domains such as medical imaging and genomics (LeCun et al., 2015). This divergence likely reflects the relatively modest sample sizes, lower feature dimensionality, and structured tabular nature of most climate-health datasets, conditions under which ensemble methods remain competitive. Nevertheless, LSTM architectures demonstrated advantages in temporal forecasting tasks, suggesting that hybrid frameworks integrating tree-based prediction with sequential modeling of lagged dynamics require systematic investigation (Schachtschneider et al., 2024).

The wide variation in predictive performance across studies, ranging from R2 values exceeding 0.90 (Jian et al., 2023; Xu et al., 2024) to approximately 0.06 (Boudreault et al., 2024), warrants contextual interpretation rather than a direct comparison. In studies targeting daily all-cause mortality deviation at the population level, the inherently low signal-to-noise ratio of the outcome constrains the achievable R2 irrespective of the algorithm choice. This is consistent with findings in which GAM achieved superior or comparable performance in such settings, reflecting its suitability for modeling smooth, non-linear exposure-response relationships in lower-complexity data structures (Boudreault et al., 2024; Wertis et al., 2023). Therefore, performance differences across studies more plausibly reflect outcome definitions, temporal continuity, spatial resolution, and data completeness than income classification or algorithm choice alone.

Building on this, a central question is whether methodological frameworks developed in data-rich HIC settings can be meaningfully transferred to LMIC contexts where climate-sensitive health burdens are the greatest (Weeda et al., 2024). The performance of RF across HIC settings, including Australia (Jian et al., 2023), South Korea (Kim and Kim, 2022), and the US (), suggests that ensemble tree-based approaches can perform robustly in high-income contexts. Similar performance has also been observed in upper-middle-income settings such as China (Wang et al., 2019; Xu et al., 2024; Wang et al., 2023), indicating that these methods may generalize beyond strictly high-income environments. In contrast, evidence from LMIC settings remains limited, with only a single study from Senegal (Toure et al., 2025) highlighting the need for further validation in data-constrained contexts. However, the absence of external validation and formal generalization studies means that transferability remains largely untested.

The performance gap observed between HIC and LMIC studies should not be interpreted as evidence of reduced algorithmic suitability in LMIC contexts. Rather, it reflects compounding deficits in both environmental monitoring and health surveillance infrastructure. On the environmental side, Africa has the least developed weather observation network globally, with station density estimated to be eight times lower than the World Meteorological Organization recommended levels (Dinku, 2019; Parker et al., 2024). Where weather stations do exist, they are often concentrated along major roads and urban centers, leaving rural areas where climate-sensitive livelihoods are most vulnerable and chronically under-monitored (Dinku, 2019). Equipment obsolescence, power supply interruptions, limited maintenance budgets, and restrictive national data policies result in frequent data gaps, discontinuities, and restricted accessibility for research use (Kaspar et al., 2021).

Consequently, the long, continuous, and high-resolution environmental time series that underpin robust ML model development in high-income settings are often unavailable for LMICs. This interpretation is further supported by vulnerability research, which demonstrates that poverty, housing quality, occupational exposure, and access to cooling mediate the relationship between environmental exposure and health outcomes (World Health Organization, 2024; Kirby et al., 2025). Illustrative evidence from South Africa highlights persistent surveillance constraints, where the absence of pathogen-resolved infectious disease data limits the granularity of predictive modeling despite substantial climate-sensitive disease burden, while diarrheal illness is typically treated symptomatically without routine pathogen identification (Dickinson et al., 2026).

4.3 Beyond heat: why flood-health prediction remains underdeveloped

The near-complete absence of flood-health prediction models represents a critical finding rather than a limitation of the search. This is supported by the breadth of the search strategy, which spanned 11 databases over a fifteen-year period and explicitly included flood-related terms as well as the well-established global significance of floods as EWE (Seneviratne et al., 2021). A single precipitation-focused study applied rainfall thresholds exceeding 150 mm/week, representing flood-adjacent conditions, to predict dengue transmission dynamics in India (Sophia et al., 2025). However, this approach modeled vector proliferation under favourable moisture conditions, where rainfall functions as a proxy for environmental conditions supporting mosquito breeding habitat availability and vector population growth, rather than direct flood-related health outcomes such as waterborne infections, injuries, or displacement-related morbidity. Rainfall showed its strongest association with dengue-related outcomes at an 11-week lag, while temperature and RH demonstrated maximum correlations at lags of 20 and 6 weeks, respectively. The study further highlighted nonlinear climate-health relationships, applying RFR to capture complex associations between environmental conditions and dengue risk (Sophia et al., 2025). This study demonstrated moderate RF performance in a vector-borne disease context, suggesting the potential applicability of ensemble tree-based methods beyond heat-related outcomes. Nevertheless, no study linked flood exposure metrics to acute flood-associated health outcomes, and broader conclusions regarding rainfall-related ML modeling remain constrained by the absence of direct algorithm comparisons and the substantial evidence gap in flood-health prediction. Consequently, the cross-EWE generalizability of RF cannot be established from the current evidence base.

This gap reflects the fundamental methodological challenges in distinguishing floods from heat exposure. Temperature exposure is continuous, ubiquitous, and reliably measured, with relatively direct physiological effects. In contrast, flood exposure is episodic, spatially fragmented, and operates through indirect pathways, including water contamination, vector proliferation, displacement, and psychosocial stress (Nusrat et al., 2022). Although, satellite-derived flood inundation mapping has improved substantially in terms of spatial resolution and near real-time availability (Notti et al., 2018; Munasinghe et al., 2018), linking exposure estimates with geocoded health records remains challenging, particularly when affected populations are displaced. Health outcome surveillance has several additional limitations. In LMIC settings, flood-associated diarrheal disease is treated syndromically without pathogen identification (Dickinson et al., 2026). In HICs, the primary flood-related burden manifests through mental health pathways, including depression, anxiety, and post-traumatic stress disorder, persisting years after the event (Dickinson et al., 2025). These outcomes are poorly captured in acute surveillance systems. These contrasting profiles underscore that flood-health prediction cannot simply adapt heat-health frameworks, but instead requires context-specific approaches to exposure characterization and outcome measurement.

4.4 Strengths, limitations, and future directions

This review benefits from the systematic application of the PECO framework, dual independent screening, and adherence to the PRISMA-2020 reporting guidelines (Page et al., 2021). By focusing on ML-based prediction rather than traditional association studies, it addresses a rapidly evolving domain with direct relevance to climate-health EWS (El Morr et al., 2024; Villanueva-Miranda et al., 2025). This review has several limitations that warrant consideration. Restrictions on English-language publications may have excluded relevant studies from China and Latin America. Heterogeneity in outcome definitions and performance metrics precluded meta-analyzes and limited direct comparability across studies. The inclusion criteria that required explicit environmental predictors may have excluded studies that used derived climate indices. The rapid evolution of ML methodologies means that recent innovations, including transformer-based and foundation models for climate applications (Camps-Valls et al., 2025), may not yet be represented in peer-reviewed literature. Furthermore, Bayesian hierarchical models and formal causal inference frameworks were not identified among the methodological approaches reported in the included studies, limiting the extent to which the current evidence base can support causal attribution or account for unobserved heterogeneity across settings. Additionally, uncertainty quantification was inconsistently reported across included studies, with few providing prediction intervals or probabilistic outputs, limiting the extent to which model outputs can inform risk communication or decision-making under uncertainty.

Risk of bias assessment using PROBAST identified unclear risk in six studies and high risk in one study (Wang et al., 2024), limiting confidence in the reliability of the reported performance estimates. The high-risk rating assigned to one study reflected absent model calibration reporting within the analysis domain rather than concerns about its exposure-outcome framework, and findings from that study were interpreted with appropriate caution throughout the synthesis. The absence of standardized performance metrics across studies further constrained interpretation, as regression and classification outcomes were evaluated using heterogeneous measures, including R2, RMSE, AUC, and accuracy. Quantitative synthesis was not performed for the all-cause mortality (n = 6) and acute heat morbidity (n = 6) subgroups because, given the metric heterogeneity noted above, direct aggregation would have been methodologically inappropriate and potentially misleading. Furthermore, in two studies, GAM achieved superior or joint-best performance relative to ML models, indicating that the comparative advantage of ML approaches is context-dependent rather than universal (Wertis et al., 2023; Boudreault et al., 2024).

Future research should prioritize systematic benchmarking of DL against ensemble methods and the development of standardized reporting frameworks, as detailed in Table 5. External validation across diverse geographic and socioeconomic settings is essential to establish generalizability. Most critically, flood-health prediction requires targeted investment, including methodological innovation in exposure characterization integrating satellite-derived inundation mapping with population mobility data, pathogen-resolved surveillance in LMIC settings to move beyond syndromic reporting, and longitudinal mental health monitoring in high-income contexts. In LMIC settings, where both environmental monitoring and health surveillance remain underdeveloped, dual investment in high-resolution sensor networks and strengthened health information systems will be essential to generate the data streams required for context-appropriate climate-health prediction.

Table 5

Review objectiveKey evidence from this reviewFuture research directions
ML modeling approachesRF was most frequently evaluated and best performing algorithm across diverse settings. DL rarely outperformed ensemble methods in direct comparisons, although LSTM showed advantages for temporal forecasting. Performance metrics varied widely, limiting cross-study comparison.Benchmark DL against ensemble methods systematically. Develop hybrid frameworks combining tree-based and sequential models. Standardize performance metric reporting. Explore integration of ML-derived climate-health predictors within hybrid mechanistic-data-driven frameworks, where climate-sensitive predictors may inform time-varying transmission rates, vector recruitment processes, or seasonal forcing parameters in infectious disease applications (Islam et al., 2024).
Environmental predictorsTemperature variables and composite heat indices (WBGT, EHF) were the strongest environmental predictors. Lagged effects frequently ranked most important. Socioeconomic variables ranked among the top predictors in three studies, contributing alongside environmental factors. LMIC studies lacked granular socioeconomic data.Standardize environmental predictor selection and lag structures. Prioritize socioeconomic indicator collection in LMICs. Integrate vulnerability data with environmental exposures.
EWE-health outcome coverageThe evidence was heat-focused (94%). One study used precipitation data to predict dengue transmission dynamics in India. No studies have linked flood metrics to acute flood-health outcomes.Extend environmental predictor research to flood-health outcomes, incorporating flood-specific variables such as inundation extent, surface runoff, and antecedent soil moisture. LMIC contexts are a priority given high flood exposure and limited predictive modeling.
Geographic coverage and generalizabilityMost studies were conducted in HICs, while LMIC representation was limited to India and Senegal. External validation was absent across all included studies, and urban populations were substantially overrepresented, limiting generalizability to rural and data-sparse settings.Prioritize LMIC model development and external validation. Invest in environmental and health data infrastructure. Apply transfer learning for resource-limited settings.

Synthesis of review objectives, key evidence, and future research directions.

5 Conclusion

This review synthesized the current evidence on the application of ML methods for predicting health outcomes associated with EWEs. Ensemble tree-based models, particularly RF, were the most frequently evaluated and best-performing algorithms across diverse geographic settings, being identified as the best-performing model in seven studies spanning the HIC and UMIC contexts. However, statistical approaches, such as GAM, demonstrated good performance in specific settings, while GB methods also performed well in certain contexts. Overall, these findings indicate that predictive accuracy depends more on outcome characteristics and data structure than on the choice of the algorithm alone. DL approaches offered limited performance advantages over ensemble methods in most included studies, with the exception of LSTM architectures, which demonstrated superiority in temporal forecasting contexts, suggesting potential for hybrid modeling frameworks. Substantial gaps remained in thematic coverage, external model validation, and geographic representation, with evidence concentrated in HICs and almost exclusively focused on heat-related exposures (n = 17, 94%). No studies have linked flood-specific environmental predictors to acute health outcomes, representing a critical gap given the global burden of flood-related morbidity. Addressing the identified gaps through methodological innovation in flood-health prediction, equitable model development in climate-vulnerable LMIC regions, and investment in environmental and health surveillance infrastructure is essential for translating predictive frameworks into operational climate-health predictive and EWS.

Statements

Data availability statement

The original contributions presented in the study are included in the article/Supplementary material, further inquiries can be directed to the corresponding author/s.

Author contributions

FA: Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Resources, Software, Validation, Visualization, Writing – original draft, Writing – review & editing. AA: Data curation, Methodology, Software, Validation, Writing – review & editing. MS: Formal analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Writing – review & editing. NR: Supervision, Validation, Writing – review & editing. MG: Supervision, Validation, Writing – review & editing. SV: Supervision, Validation, Writing – review & editing. DN: Supervision, Validation, Writing – review & editing. ND: Formal analysis, Methodology, Validation, Writing – review & editing. LS: Formal analysis, Methodology, Validation, Writing – review & editing. ML: Funding acquisition, Project administration, Supervision, Writing – review & editing. FH-M: Methodology, Validation, Writing – review & editing. OM: Methodology, Validation, Writing – review & editing. SN: Funding acquisition, Project administration, Validation, Writing – review & editing. NN-R: Project administration, Resources, Validation, Writing – review & editing.

Funding

The author(s) declared that financial support was received for this work and/or its publication. This work was funded by the Global Health Transformation Program of the National Institute for Health and Care Research (NIHR) through the WEATHER project (Grant No. NIHR204825).

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

The author MS declared that they were an editorial board member of Frontiers, at the time of submission. This had no impact on the peer review process and the final decision.

Generative AI statement

The author(s) declared that Generative AI was used in the creation of this manuscript. The author(s) declare that Gen AI tools were used for language editing. Grammarly and QuillBot were used to proofread and improve the manuscript's clarity.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fclim.2026.1843933/full#supplementary-material

References

  • 1

    AdjeiF. (2025). Artificial intelligence and machine learning in environmental health science: a review of emerging applications. SSRN Electr. J. doi: 10.2139/ssrn.5361829

  • 2

    AndersonG. B.OlesonK. W.JonesB.PengR. D. (2018). Classifying heatwaves: developing health-based models to predict high-mortality versus moderate United States heatwaves. Clim. Change146, 439453. doi: 10.1007/s10584-016-1776-0

  • 3

    AsaagaF. A.TomudeE. S.RickardsN. J.HassallR.SarkarS.PurseB. V. (2024). Informing climate-health adaptation options through mapping the needs and potential for integrated climate-driven early warning forecasting systems in South Asia–a scoping review. PLoS ONE19:e0309757. doi: 10.1371/journal.pone.0309757

  • 4

    AsciertoA.Ouambo TallaA. W.SanvidoA.SeveriP.TrabuioA.CeccoliniF.et al. (2025). The heat, the heart and beyond: a narrative review of the many ways climate change impacts human health. Front. Clim. 7:1647942. doi: 10.3389/fclim.2025.1647942

  • 5

    AyazF.AshfordA.DickinsonN.SpencerL. H.LynchM.ShakirM. Z. (2025). “Environmental drivers and machine learning models for climate-health risk prediction from extreme weather events: a global systematic review,” in Technical Report CRD420251077655 (PROSPERO). Available online at: https://www.crd.york.ac.uk/PROSPERO/view/CRD420251077655 (Accessed July 03, 2025).

  • 6

    BallesterJ.LoweR.DiggleP. J.RodóX. (2016). Seasonal forecasting and health impact models: challenges and opportunities. Ann. N. Y. Acad. Sci. 1382, 820. doi: 10.1111/nyas.13129

  • 7

    BartlowA. W.ManoreC.XuC.KaufeldK. A.Del ValleS.ZiemannA.et al. (2019). Forecasting zoonotic infectious disease response to climate change: mosquito vectors and a changing environment. Vet. Sci. 6:40. doi: 10.3390/vetsci6020040

  • 8

    BermanJ. D.EbisuK.PengR. D.DominiciF.BellM. L. (2017). Drought and the risk of hospital admissions and mortality in older adults in western USA from 2000 to 2013: a retrospective study. Lancet Planet. Health1, e17e25. doi: 10.1016/S2542-5196(17)30002-5

  • 9

    Berrang-FordL.SietsmaA. J.CallaghanM.MinxJ. C.ScheelbeekP. F.HaddawayN. R.et al. (2021). Systematic mapping of global research on climate and health: a machine learning review. Lancet Planet. Health5, e514525. doi: 10.1016/S2542-5196(21)00179-0

  • 10

    BoudreaultJ.CampagnaC.ChebanaF. (2023). Machine and deep learning for modelling heat-health relationships. Sci. Total Environ. 892:164660. doi: 10.1016/j.scitotenv.2023.164660

  • 11

    BoudreaultJ.CampagnaC.ChebanaF. (2024). Revisiting the importance of temperature, weather and air pollution variables in heat-mortality relationships with machine learning. Environ. Sci. Pollut. Res. 31, 1405914070. doi: 10.1007/s11356-024-31969-z

  • 12

    BoudreaultJ.LamotheF.CampagnaC.ChebanaF. (2025). Machine learning for modelling the health impacts of extreme heat: a comprehensive literature review. Environ. Int. 206:109965. doi: 10.1016/j.envint.2025.109965

  • 13

    Camps-VallsG.Fernández-TorresM. .-,Á CohrsK.-H.HöhlA.CastellettiA.et al. (2025). Artificial intelligence for modeling and understanding extreme weather and climate events. Nat. Commun. 16:1919. doi: 10.1038/s41467-025-56573-8

  • 14

    Ciecierski-HolmesT.SinghR.AxtM.BrennerS.BarteitS. (2022). Artificial intelligence for strengthening healthcare systems in low-and middle-income countries: a systematic scoping review. NPJ Digit. Med. 5:162. doi: 10.1038/s41746-022-00700-y

  • 15

    CôtéJ.-N.GermainM.LevacE.LavigneE. (2024). Vulnerability assessment of heat waves within a risk framework using artificial intelligence. Sci. Total Environ. 912:169355. doi: 10.1016/j.scitotenv.2023.169355

  • 16

    Darsha JayaminiW. K.MirzaF.Asif NaeemM.ChanA. H. Y. (2024). Investigating machine learning techniques for predicting risk of asthma exacerbations: a systematic review. J. Med. Syst. 48:49. doi: 10.1007/s10916-024-02061-3

  • 17

    DickinsonN.SpencerL. H.MillerC.Nadesan-ReddyN.ViririS.ShakirM. Z.et al. (2026). A systematic review investigating emerging trends between extreme weather events (ewes) and infectious disease outbreaks in South Africa. Front. Public Health14:1778784. doi: 10.3389/fpubh.2026.1778784

  • 18

    DickinsonN.SpencerL. H.YangS.MillerC.HursthouseA.LynchM. (2025). Extreme weather events in the UK and resulting public health outcomes. Int. J. Public Health70:1607904. doi: 10.3389/ijph.2025.1607904

  • 19

    DinkuT. (2019). “Challenges with availability and quality of climate data in Africa,” in Extreme Hydrology and Climate Variability, eds. MelesseA. M.AbtewW.SenayG. (Amsterdam: Elsevier), 7180. doi: 10.1016/B978-0-12-815998-9.00007-5

  • 20

    El MorrC.OzdemirD.AsdaahY.SaabA.El-LahibY.SokhnE. S. (2024). AI-based epidemic and pandemic early warning systems: a systematic scoping review. Health Inf. J. 30:14604582241275844. doi: 10.1177/14604582241275844

  • 21

    FerrariG. N.LealG. C. L.OssaniP. C.GaldamezE. V. C. (2025). Investigation of the usage of machine learning to explore the impacts of climate change on occupational health: a systematic review and research agenda. Front. Public Health13:1578558. doi: 10.3389/fpubh.2025.1578558

  • 22

    FranchiniM.MannucciP. M. (2015). Impact on human health of climate changes. Eur. J. Intern. Med. 26, 15. doi: 10.1016/j.ejim.2014.12.008

  • 23

    ImC.KimW.KimH. (2025). Explainable machine learning for heat-related illness prediction: an XGBoost-SHAP approach using Korean meteorological data. Bioengineering12:1276. doi: 10.3390/bioengineering12111276

  • 24

    Intergovernmental Panel on Climate Change (IPCC) (2023). “Sixth assessment report (ar6),”Technical Report (Intergovernmental Panel on Climate Change (IPCC)). Available online at: https://www.ipcc.ch/assessment-report/ar6/ (Accessed March 16, 2026).

  • 25

    IslamM. S.ShahrearP.SahaG.AtaullhaM.RahmanM. S. (2024). Mathematical analysis and prediction of future outbreak of dengue on time-varying contact rate using machine learning approach. Comput. Biol. Med. 178:108707. doi: 10.1016/j.compbiomed.2024.108707

  • 26

    Jacques-DumasV.RagoneF.BorgnatP.AbryP.BouchetF. (2022). Deep learning-based extreme heatwave forecast. Front. Clim. 4:789641. doi: 10.3389/fclim.2022.789641

  • 27

    JianL.PatelD.XiaoJ.JanszJ.YunG.LinT.et al. (2023). Can we use a machine learning approach to predict the impact of heatwaves on emergency department attendance?Environ. Res. Commun. 5:045005. doi: 10.1088/2515-7620/acca6e

  • 28

    JordanM. I.MitchellT. M. (2015). Machine learning: trends, perspectives, and prospects. Science349, 255260. doi: 10.1126/science.aaa8415

  • 29

    KafiK. M.PonrahonoZ. (2024). Advances in weather and climate extreme studies: a systematic comparative review. Discov. Geosci. 2:66. doi: 10.1007/s44288-024-00079-1

  • 30

    KasparF.AnderssonA.ZieseM.HollmannR. (2021). Contributions to the improvement of climate data availability and quality for sub-saharan Africa. Front. Clim. 3:815043. doi: 10.3389/fclim.2021.815043

  • 31

    KaushikA.BarcellonaC.MandyamN. K.TanS. Y.TrompJ. (2025). Challenges and opportunities for data sharing related to artificial intelligence tools in health care in low-and middle-income countries: systematic review and case study from Thailand. J. Med. Internet Res. 27:e58338. doi: 10.2196/58338

  • 32

    KeD.TakahashiK.TakakuraJ.TakaraK.KamranzadB. (2023). Effects of heatwave features on machine-learning-based heat-related ambulance calls prediction models in Japan. Sci. Total Environ. 873:162283. doi: 10.1016/j.scitotenv.2023.162283

  • 33

    KhatibuS.NgowiD. E. E. (2025). Effectiveness of climate information services in sub-saharan Africa's agricultural sector: a systematic review of what works, what doesn't work, and why. Front. Clim. 7:1616691. doi: 10.3389/fclim.2025.1616691

  • 34

    KimY.KimY. (2022). Explainable heat-related mortality with random forest and shapley additive explanations (SHAP) models. Sustain. Cities Soc. 79:103677. doi: 10.1016/j.scs.2022.103677

  • 35

    KirbyN. V.TetzlaffE. J.KiddS. A.BrownE. E.BezgrebelnaM.YoonL.et al. (2025). Susceptibility of persons with schizophrenia to extreme heat: a critical review of physiological, behavioural, and social factors. Sci. Total Environ. 995:179965. doi: 10.1016/j.scitotenv.2025.179965

  • 36

    KitchenhamB. (2004). Procedures for Performing Systematic Reviews. Keele: Keele University.

  • 37

    LawrenceW. R.SoimA.ZhangW.LinZ.LuY.LiptonE. A.et al. (2021). A population-based case-control study of the association between weather-related extreme heat events and low birthweight. J. Dev. Origins Health Dis. 12, 335342. doi: 10.1017/S2040174420000392

  • 38

    LeCunY.BengioY.HintonG. (2015). Deep learning. Nature521, 436444. doi: 10.1038/nature14539

  • 39

    LiuT.KrentzA.LuL.CurcinV. (2025). Machine learning based prediction models for cardiovascular disease risk using electronic health records data: systematic review and meta-analysis. Eur. Heart J. Digit. Health6, 722. doi: 10.1093/ehjdh/ztae080

  • 40

    MachariaD.MugaboL.KasitiF.NoriegaA.MacDonaldL.ThomasE. (2023). Streamflow and flood prediction in Rwanda using machine learning and remote sensing in support of rural first-mile transport connectivity. Front. Clim. 5:1158186. doi: 10.3389/fclim.2023.1158186

  • 41

    McGowanJ.SampsonM.SalzwedelD. M.CogoE.FoersterV.LefebvreC. (2016). Press peer review of electronic search strategies: 2015 guideline statement. J. Clin. Epidemiol. 75, 4046. doi: 10.1016/j.jclinepi.2016.01.021

  • 42

    MorganR. L.WhaleyP.ThayerK. A.SchünemannH. J. (2018). Identifying the PECO: a framework for formulating good questions to explore the association of environmental and other exposures with health outcomes. Environ. Int. 121, 10271031. doi: 10.1016/j.envint.2018.07.015

  • 43

    MunasingheD.CohenS.HuangY.-F.TsangY.-P.ZhangJ.FangZ. (2018). Intercomparison of satellite remote sensing-based flood inundation mapping techniques. J. Am. Water Resour. Assoc. 54, 834846. doi: 10.1111/1752-1688.12626

  • 44

    NishimuraT.RashedE. A.KoderaS.ShirakamiH.KawaguchiR.WatanabeK.et al. (2021). Social implementation and intervention with estimated morbidity of heat-related illnesses from weather data: a case study from Nagoya city, Japan. Sustain. Cities Soc. 74:103203. doi: 10.1016/j.scs.2021.103203

  • 45

    NottiD.GiordanD.CalóF.PepeA.ZuccaF.GalveJ. P. (2018). Potential and limitations of open satellite data for flood mapping. Rem. Sens. 10:1673. doi: 10.3390/rs10111673

  • 46

    NusratF.HaqueM.RollendD.ChristieG.AkandaA. S. (2022). A high-resolution earth observations and machine learning-based approach to forecast waterborne disease risk in post-disaster settings. Climate10:48. doi: 10.3390/cli10040048

  • 47

    OhashiY.IharaT.OkaK.TakaneY.KikegawaY. (2023). Machine learning analysis and risk prediction of weather-sensitive mortality related to cardiovascular disease during summer in Tokyo, Japan. Sci. Rep. 13:17020. doi: 10.1038/s41598-023-44181-9

  • 48

    OuzzaniM.HammadyH.FedorowiczZ.ElmagarmidA. (2016). Rayyan-a web and mobile app for systematic reviews. Syst. Rev. 5:210. doi: 10.1186/s13643-016-0384-4

  • 49

    PageM. J.McKenzieJ. E.BossuytP. M.BoutronI.HoffmannT. C.MulrowC. D.et al. (2021). The prisma 2020 statement: an updated guideline for reporting systematic reviews. BMJ372:n71. doi: 10.1136/bmj.n71

  • 50

    ParkerD. J.BainC. L.ChemelC.et al. (2024). Challenges and ways forward for sustainable weather and climate services in Africa. Nat. Commun. 15:2713. doi: 10.1038/s41467-024-46742-6

  • 51

    RochaJ.OliveiraS.VianaC. M.RibeiroA. I. (2022). “Climate change and its impacts on health, environment and economy,” in One Health, eds. PrataJ. C.Isabel RibeiroA.Rocha-SantosT. (Amsterdam: Elsevier), 253279. doi: 10.1016/B978-0-12-822794-7.00009-5

  • 52

    SchachtschneiderR.Saynisch-WagnerJ.Sánchez-BenitezA.ThomasM. (2024). Neural network based estimates of the climate impact on mortality in germany: application to storyline climate simulations. Sci. Rep. 14:26074. doi: 10.1038/s41598-024-77398-3

  • 53

    SeneviratneS. I.ZhangX.AdnanM.BadiW.DereczynskiC.Di LucaA.et al. (2021). “Weather and climate extreme events in a changing climate,” in Climate Change 2021: The Physical Science Basis. Contribution of Working Group I to the Sixth Assessment Report of the Intergovernmental Panel on Climate Change, eds. Masson-DelmotteV.ZhaiP.PiraniA.ConnorsS. L.PanC.BergerS.CaudN.ChenY.GoldfarbL.GomisM. I.HuangM.LeitzellK.LonnoyE.MatthewsJ. B. R.MaycockT. K.WaterfieldT.YelekiO.YuR.ZhouB. (Cambridge: Cambridge University Press), 15131766.

  • 54

    SophiaY.RoxyM. K.MurtuguddeR.KaripotA.SapkotaA.DasguptaP.et al. (2025). Dengue dynamics, predictions, and future increase under changing monsoon climate in india. Sci. Rep. 15:1637. doi: 10.1038/s41598-025-85437-w

  • 55

    SsebyalaS. N.KintuT. M.MuganziD. J.DresserC.DemetresM. R.LaiY.et al. (2024). Use of machine learning tools to predict health risks from climate-sensitive extreme weather events: a scoping review. PLoS Clim. 3:e0000338. doi: 10.1371/journal.pclm.0000338

  • 56

    TeshaleA. B.HtunH. L.VeredM.OwenA. J.Freak-PoliR. (2024). A systematic review of artificial intelligence models for time-to-event outcome applied in cardiovascular disease risk prediction. J. Med. Syst. 48:68. doi: 10.1007/s10916-024-02087-7

  • 57

    ToureM.SyI.DioufI.GueyeO.BekeleE.BhuiyanM. A. E.et al. (2025). Machine learning-based prediction of heatwave-related hospitalizations: a case study in Matam, Senegal. Int. J. Environ. Res. Public Health22:1349. doi: 10.3390/ijerph22091349

  • 58

    Villanueva-MirandaI.XiaoG.XieY. (2025). Artificial intelligence in early warning systems for infectious disease surveillance: a systematic review. Front. Public Health13:1609615. doi: 10.3389/fpubh.2025.1609615

  • 59

    WangJ.NikolaouN.an der HeidenM.IrrgangC. (2024). High-resolution modeling and projection of heat-related mortality in Germany under climate change. Commun. Med. 4:206. doi: 10.1038/s43856-024-00643-3

  • 60

    WangS.CaiW.TaoY.SunQ. C.WongP. P. Y.ThongkingW.et al. (2023). Nexus of heat-vulnerable chronic diseases and heatwave mediated through tri-environmental interactions: a nationwide fine-grained study in Australia. J. Environ. Manag. 325:116663. doi: 10.1016/j.jenvman.2022.116663

  • 61

    WangY.BuL.WeiZ.ChengY.FengL.WangS. (2025). Extreme urban temperature exposure and preterm birth: spatial-temporal risk zone prediction using machine learning models. Environ. Res. 284:122230. doi: 10.1016/j.envres.2025.122230

  • 62

    WangY.SongQ.DuY.WangJ.ZhouJ.DuZ.et al. (2019). A random forest model to predict heatstroke occurrence for heatwave in China. Sci. Total Environ. 650, 30483053. doi: 10.1016/j.scitotenv.2018.09.369

  • 63

    WeedaL. J.BradshawC. J.JudgeM. A.SaraswatiC. M.Le SouëfP. N. (2024). How climate change degrades child health: a systematic review and meta-analysis. Sci. Total Environ. 920:170944. doi: 10.1016/j.scitotenv.2024.170944

  • 64

    WertisL.SuggM. M.RunkleJ. D.RaoD. (2023). Socio-environmental determinants of mental and behavioral disorders in youth: a machine learning approach. Geohealth7:e2023GH000839. doi: 10.1029/2023GH000839

  • 65

    WolffR. F.MoonsK. G.RileyR. D.WhitingP. F.WestwoodM.CollinsG. S.et al. (2019). Probast: a tool to assess the risk of bias and applicability of prediction model studies. Ann. Intern. Med. 170, 5158. doi: 10.7326/M18-1376

  • 66

    World Bank (2024). “World bank country classifications by income level for 2024-2025,” in Technical Report (World Bank). Available online at: https://blogs.worldbank.org/en/opendata/world-bank-country-classifications-by-income-level-for-2024-2025 (Accessed July 03, 2025).

  • 67

    World Health Organization (2024). “Heat and health,” in Technical Report (World Health Organization). Available online at: https://www.who.int/news-room/fact-sheets/detail/climate-change-heat-and-health (Accessed November 17, 2025).

  • 68

    XiY.WettsteinZ. S.KshirsagarA. V.LiuY.ZhangD.HangY.et al. (2024). Elevated ambient temperature associated with increased cardiovascular disease-risk among patients on hemodialysis. Kidney Int. Rep. 9, 29462955. doi: 10.1016/j.ekir.2024.07.015

  • 69

    XuH.GuoS.ShiX.WuY.PanJ.GaoH.et al. (2024). Machine learning-based analysis and prediction of meteorological factors and urban heatstroke diseases. Front. Public Health12:1420608. doi: 10.3389/fpubh.2024.1420608

  • 70

    YenewC.BayehG. M.GebeyehuA. A.EnawgawA. S.AsmareZ. A.EjiguA. G.et al. (2025). Scoping review on assessing climate-sensitive health risks. BMC Public Health25:914. doi: 10.1186/s12889-025-22148-x

Summary

Keywords

climate-sensitive health outcome prediction, environmental exposure, extreme weather events, machine learning, systematic review

Citation

Ayaz F, Ashford A, Shakir MZ, Ramzan N, Gebreslasie M, Viriri S, Ndzi D, Dickinson N, Spencer LH, Lynch M, Henriquez-Mui F, Mahomed O, Naidoo S and Nadesan-Reddy N (2026) Environmental predictors and machine learning models for climate-health risk prediction from extreme weather events: a global systematic review. Front. Clim. 8:1843933. doi: 10.3389/fclim.2026.1843933

Received

31 March 2026

Revised

25 June 2026

Accepted

30 June 2026

Published

21 July 2026

Volume

8 - 2026

Edited by

Nir Krakauer, City College of New York (CUNY), United States

Reviewed by

Chaeyeong Im, Ministry of National Defense, Republic of Korea

Md Shahidul Islam, University of Tennessee at Chattanooga, United States

Updates

Copyright

*Correspondence: Fahad Ayaz,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics