Abstract
Introduction:
Flooding poses an escalating threat to Kisumu County, Kenya, driven by climate change and rapid urbanization.
Methods:
This study develops an integrated framework combining remote sensing, GIS, and machine learning for flood risk mapping and forecasting using multi‐source geospatial data (Sentinel‐1 SAR, Landsat 8, SRTM DEM, SMAP soil moisture) and meteorological data (2014‐2025). Four machine learning models random forest (RF), artificial neural network (ANN), deep neural network (DNN), and convolutional neural network (CNN) were developed and validated.
Results:
The ANN achieved the highest accuracy (R2 = 0.910, RMSE = 0.038). Excluding RF (R2 = 0.520) from the ensemble improved performance to R2 = 0.916 (RMSE = 0.035), a 22% error reduction. Flood risk mapping classified Kisumu County into five categories, revealing that 40.9% (857.8 km2) faces moderate to very high risk, with critical hotspots along the Lake Victoria shoreline and in Kisumu City and Ahero. Slope (20%), NDBI (15%), soil moisture (15%), and stream proximity (15%) were identified as dominant flood drivers. Uncertainty quantification revealed low model variance (R2 σ = 0.006) and demonstrated that aleatoric uncertainty (62%) dominates epistemic uncertainty (38%).
Discussion/Conclusion:
This framework provides a replicable, data‐driven methodology for flood risk assessment in data‐scarce regions globally.
1 Introduction
Floodplains have long been attractive for settlement due to their fertile soils, supporting diverse plant and animal life as well as agricultural communities (). However, flooding poses significant risks to both natural ecosystems and human settlements. The physical manifestation of flooding varies depending on water flow rates, as well as the geometry and orientation of the affected area (). Future flood exposure is expected to increase due to multiple factors, including sea-level rise, coastal urbanization, deforestation, deteriorating infrastructure, and rural-urban migration (Moftakhari et al., 2018).
Climate change has also introduced water scarcity and uncharacteristic flooding in traditionally humid temperate zones (Oleku, 2024). Projections indicate increased precipitation in the Lake Victoria Basin, which will likely worsen flooding in Kisumu and surrounding sub-basins (Ombogo, 2016). For instance, the Sondu Miriu River Basin has observed an 18.27 mm increase in annual rainfall, with seasonal increases in all periods except June–September (Ochieng’-Springer, 2022). Flooding in Kisumu County is increasingly magnified by climate change, which alters rainfall patterns, intensity, and frequency while heightening extreme weather conditions (Oluchiri, 2025). A 30-year climate trend analysis reveals significant fluctuations in rainfall and temperature, underscoring the heightened vulnerability of residents of informal settlements to climate-related disasters (Masimbe, 2018).
Globally, floods have caused severe devastation in recent decades, surpassing other natural hazards in frequency, human casualties, economic damage, and destruction of arable land and housing. Of more than 7,000 recorded natural disasters, nearly 75% are water-related, with floods being the most prevalent, constituting about one-third of these events (). The impacts of flooding extend beyond immediate losses, with long-term consequences for livelihoods and property values. Housing prices typically decline post-flood but recover after 2–3 years, although substantial outmigration can disrupt local investment (Masese et al., 2016).
Historical data highlight the severity of past floods in Kisumu County, with more than 90% of Nyamasaria residents and 57%–88% of Manyatta tenants reporting flood exposure (Laji and Ayonga, 2024). Many cite the 1997 EL Niño and 2008–2009 floods as the most devastating. Such trends align with global patterns, where radiative forcing intensifies rainfall extremes (Njeru, 2025).
2 Statement of the problem
A major challenge lies in the scarcity and inconsistency of streamflow records, which are essential for calibrating statistical and hydrological models. These records may be incomplete, short-term, or affected by measurement errors, inhomogeneities, or restricted accessibility (). Additionally, insufficient spatial data and limited knowledge of factors influencing streamflow further hinder model accuracy.
Physical hydrological models, on the other hand, demand extensive expertise in hydrologic processes and rely heavily on high-quality in situ data, restricting their applicability to well-monitored regions (). Effective flood mapping requires continuous hydroclimatic monitoring to identify early disaster indicators, yet many systems struggle with data gaps (Mehmood and Rasmy, 2020). Urban flood detection presents additional challenges, as image analysis is complicated by unimodal or indistinct bimodal histogram distributions, making threshold determination difficult (Sadiq et al., 2022).
3 Literature review
3.1 Role of geospatial information science and remote sensing flood modeling
Remote sensing enhances flood modeling by providing key variables such as soil moisture, with systematic errors reduced through batch calibration and random errors addressed through data assimilation (Li, 2016). When calibrated with hydraulic models, flood maps become powerful tools in damage assessment and emergency response (Sharma et al., 2019). Remote sensing has increased data availability for hydrological modeling, mostly in data-scarce regions. Landsat imagery at 30 m resolution and Synthetic Aperture Radar (SAR) imagery at 30 m and 10 m resolution are widely used in flood mapping (; Li et al., 2018). With advances in SAR processing, building-shadow removal and region-growing algorithms are effective in urban flooding (Tanim et al., 2022). For example, SAR was successfully applied to identify 75% of inundated areas, with waterline positions accurate within 20 m of ground data (Mason et al., 2009), making remote sensing (RS) a very valuable tool to obtain historical and present images and forecast flooding.
Remote sensing techniques provide less expensive and faster options for accessing spatial data about a flood, especially in physically inaccessible areas (Opolot, 2013). A digital elevation model is the most inexpensive and efficient method of estimating flood depth from hydrological and remotely sensed data (Sheikh Kumran et al., 2020). The Advanced Microwave Scanning Radiometer (AMSR-E) for NASA’s Earth Observing System, NASA’s Soil Moisture Active and Passive (SMAP) mission, and the European Space Agency (ESA) Soil Moisture and Ocean Salinity (SMOS) mission all provide global soil moisture datasets (Wang, 2018). They have been used effectively due to their short temporal resolution (Md. Shahinoor Rahman, 2017). Satellite data are also available from the Landsat multi-spectral scanner (MSS) with 80 m spatial resolution since early 1972. Before then, single broad-spectrum aerial photographs were used for mapping of hydrogeological units and geomorphological features ().
Geospatial techniques, on the other hand, facilitate hydrological models by simplifying data collection, analysis, interpretation, and presentation of information (Opolot, 2013). Geospatial tools are very versatile, especially in spatial analysis, modeling, visualization, data processing, and management (Wang, 2018). Geospatial software has emerged to include more hydrology-specific tools, increasing the popularity of this level of integration and providing environments for generating flood probability maps (). GIS and forecasting models are powerful instruments to support decision makers in identifying risk areas. The integration of models with GIS can produce results rapidly, which is effective for project implementation ().
Geospatial and RS-based flood maps derived from elevation, water depth, and land use data have informed urban resilience planning (Sengupta, 2024). These applications align with the Sendai Framework goals of disaster risk reduction and resilience building (Munawar et al., 2022). Remote sensing and GIS have also proved resourceful in flood management, as forecasted areas of potential flood risk could be produced by using the overlaying function of a GIS to combine land cover maps with the flood forecast zones (Opolot, 2013). The flood risk map is one of the unique and mandatory outputs of risk zoning maps provided by GIS and RS by mixing different layers of data (Sheikh Kumran et al., 2020). Remote sensing and GIS have also been effective in time series analysis of the dynamic evolution of flood risk research hotspots (). Remote sensing and GIS have been widely used to study natural disasters, specifically floods and flood susceptibility mapping (FSM), with the help of RS and GIS, using different models such as neural networks and random forests (RFs) to compare the prediction power of the data mining algorithms (). For effective monitoring of the environmental aspects, geospatial technologies, including RS and GIS, are effective and useful (Paul et al., 2020).
3.2 Climate modeling
Meteorological inputs, particularly rainfall, remain central to flood forecasting (). Rainfall intensity, duration, and stream network characteristics are critical factors influencing flood forecasting (Puttinaovarat and Horkaew, 2020). Numerical weather prediction (NWP) models provide essential forecast fields such as rainfall, temperature, and potential evapotranspiration, often at fine spatial (0.09° × 0.09°) and temporal (3 h) resolutions ().
For real-time monitoring, satellite-based rainfall products enhance accuracy by combining thermal infrared (TIR) and passive microwave (PM) data. Widely applied approaches include Precipitation Estimation from Remote Sensing Information using Artificial Neural Networks (PERSIANN), the Climate Prediction Center MORPHing (CMORPH) technique, and Global Satellite Mapping of Precipitation (GSMaP) (Veerakachen and Raksapatcharawong, 2015). The NASA Tropical Rainfall Measuring Mission (TRMM) Multisatellite Precipitation Analysis provides quasi-global coverage (50 S–50 N) at 0.25° resolution, calibrated with radar and microwave sensors (Wu et al., 2012).
Accurate flood forecasting depends on reliable rainfall estimation, which drives hydrological models that forecast streamflow (Shi et al., 2015). In semi-arid regions, the absence of rain gauges significantly reduces model accuracy, as even advanced rainfall-runoff models struggle with incomplete datasets (Nguyen et al., 2021). Neural networks have also been applied in rainfall forecasting, demonstrating strong performance in hydrological applications ().
Satellite-derived rainfall offers a viable alternative, although validation for the study region is essential (). High-resolution datasets such as Climate Hazards Group Infrared Precipitation with Stations (CHIRPS) provide quasi-global coverage at daily, pentadal, and monthly scales, enabling the analysis of both current and historical precipitation patterns (Wanzala, 2022). Together, these advances show that the integration of physical, statistical, and machine learning (ML) models with high-quality climatic data is essential for reliable flood forecasting.
3.3 Hydrology modeling
Integrated hydrological modeling, combining meteorological and morphological inputs, has proven effective in assessing and forecasting flood risk mapping. For instance, areas near a stream mouth are forecasted to experience greater flood events (Shen et al., 2024). The use of multiple hydrological models with different structures helps quantify structural uncertainties, thereby enhancing forecast reliability (Zhou et al., 2021).
Operational flood forecasting systems (FFSs) integrate hydrological models, observational data, and user-friendly interfaces to provide timely forecasting that enhances preparedness and reduces losses. Conceptual frameworks of large-scale FFSs identify the essential components and outputs for effective operation (Wanzala et al., 2025). Probabilistic river flow forecasting typically uses ensemble rainfall forecasts to drive hydrological models (Wu et al., 2020a). Parameter optimization is central to improving forecast performance as it selects the best-fit model parameters from historical datasets. Because floods are complex, variable, and partly stochastic, techniques such as clustering analysis of similar flood types can improve parameter tuning and enhance forecast accuracy (Tang et al., 2023).
Flood susceptibility mapping (FSM) also combines statistical and hydrological models with emerging ML techniques, such as ANN and RF, to model complex flood dynamics more accurately (Munyi, 2024). Flood mapping and forecasting on continuous hydroclimatic data collection to assess watershed conditions and identify risks (Mehmood and Rasmy, 2020).
3.4 Topographic modeling
The vertical accuracy of terrain datasets strongly affects flood model forecasting, more than horizontal resolution, especially in areas of extreme slope variation (Peramuna et al., 2025). Therefore, high-resolution topographic data, particularly from LiDAR, is valuable for flood modeling, enabling explicit representation of small floodplain features within digital elevation models (DEMs) (Thomas Steven Savage et al., 2016). Model sensitivity also depends on DEM resolution and roughness values: higher-resolution DEMs generally improve performance, although excessively fine datasets are not always essential when combined with appropriate roughness layers and computational point spacings (Yalcin and Akyurek, 2020; Moghim et al., 2023). Advances in DEM development, including algorithmic corrections, have produced datasets that more accurately simulate inundation (Shastry and Durand, 2019).
Improved floodplain topography is especially important for high-risk regions, such as the lower Zambezi River Basin in Mozambique, where enhanced DEMs could significantly improve hazard forecasting and management (Vilanculos, 2015). Case studies in U.S. cities such as New York and Baltimore further demonstrate the utility of topographic indices, such as the topographic wetness index (TWI) and sink depth, in identifying flood-prone locations. Nuisance flood reports often align with areas of deep sinks and high TWI, particularly when maximum index values are considered around reported sites ().
Integrating geomorphological and hydrological methods improves flood mapping accuracy for areas with frequent recurrence intervals (Lastra et al., 2008). Geospatial-based geomorphological analysis, therefore, supports flood risk management by providing spatially explicit insights for policymakers (Tsanakas et al., 2025).
3.5 Anthropogenic modeling
Advanced hydrological models increasingly account for human influences on flood dynamics. For instance, the Grid Xin’anjiang-Haihe (GXH) model explicitly incorporates groundwater overexploitation and reservoir operations by modifying runoff generation and concentration algorithms, thereby improving forecasting of flood peaks and timing (). Similarly, GIS-based analytical hierarchy process (AHP) and flood hazard potential index approaches have been used in Iran’s Gorganrood River Basin to prioritize sub-watersheds by integrating both natural and anthropogenic drivers (Rahmati et al., 2016).
Techniques such as AHP combined with GIS allow for the separate and integrated assessment of natural versus human-induced drivers. These models assign weights to land use, urban development, and infrastructure, and validate forecasting against historical flood records. Evidence suggests that models incorporating anthropogenic factors alone can perform as well as, or better than, those using natural variables, achieving area under the ROC curve values up to 79.5% (Rahmati et al., 2016; Stefanidis and Stathis, 2013).
3.6 Flood forecasting
Flood forecasting remains a complex challenge due to uncertainties in climatic and environmental factors. Enhancing accuracy requires advanced technologies and automated forecasting systems. In the past 20 years, RS has played a critical role in flood forecasting, especially during the pre-disaster phase of disaster management (Munawar et al., 2022).
Machine learning in flood forecasting offers efficient alternatives to computationally intensive hydraulic models. ML-based systems integrate historical records, real-time data, and predictive analytics to support decision-making in reservoir and dam operations (). Recent studies highlight ML’s ability to improve flood inundation modeling, addressing the limitations of traditional ML-based approaches (Nevo et al., 2022). Algorithms such as ANNs and classical learning techniques such as support vector machines (SVMs) are widely applied in rainfall estimation, flood forecasting, and reservoir inflow prediction (Nguyen and Chen, 2020). Similarly, fuzzy logic models have been successfully used to predict river levels and flood discharge (Nguyen and Chen, 2020).
Despite advances, flood forecasting in many developing nations still lacks access to cutting-edge technologies. Reliable forecasting depends on robust predictive models validated against historical datasets (Lawal et al., 2021). Hybrid methods combining both physical models with ML-based error correction show promise. For example, Bayesian linear regression has reduced forecasting errors by up to 44.4% compared to traditional methods such as MIKE 11 (Noymanee and Theeramunkong, 2019).
Forecast outputs are often categorized into severity levels: no flood, flooding below 20 cm, moderate flooding of 20–49 cm, and severe flooding exceeding 50 cm (Puttinaovarat and Horkaew, 2020). Innovative sensor networks and statistical models have been developed to improve real-time forecasting. Regression-based approaches, for instance, outperform conventional hydrological methods for short-term 1-h forecasting and perform competitively for longer 24-h forecasts ().
3.7 Flood risk
Risk is the probability of disaster occurrence and the potential degree of loss, shaped by the interplay of hazard, exposure, and vulnerability (). Flood risk assessment integrates current hydrological conditions and climate projections to predict future river discharge, thereby aiding in flood mitigation strategies (Sharma et al., 2019). Geographic variations in vulnerability, building age, and structural value further refine flood risk assessments, enhancing their validity (Wing et al., 2020).
Risk analysis employs GIS to generate hazards and risk maps, combining vulnerability surveys with topographic data. These maps identify critical weaknesses in flood defense systems, such as dike failures, seepage, or drainage infiltration, as evidenced by the 1997 Oder flood (Plate, 2002). Effective flood risk management balances existing risk mitigation with long-term planning, incorporating both structural and non-structural measures.
The European Union (EU) Floods Directive outlines a three-step process: preliminary floodplain delineation, hazard mapping, and risk assessment. Hazard maps depict maximum inundation depths, while risk maps estimate monetary losses per grid cell (Tsakiris, 2014). Future flood risks will be driven by climatic shifts, land-use changes, and socioeconomic exposure (Kundzewicz et al., 2019). Global assessments highlight disparities in flood protection, with high-income regions often shielded against extreme events compared to lower-income areas (Winsemius et al., 2016).
Human memory of disasters is short-lived. Risk awareness peaks post-event but fades over time, leading to cyclical underestimation (). Damage estimation involves classifying at-risk elements, assessing exposure, and evaluating susceptibility (). Risk is therefore quantified as the product of disaster likelihood and societal vulnerability ().
Flood risk mapping spans from global to local scales. While global analyses, such as European flood projections or natural hazard atlases, provide broad insights, local-scale maps (1:2,000–1:20,000) are critical for site-specific defenses (Merz et al., 2007). National assessments, such as the UK’s Risk Assessment for Flood and Coastal Defense Systems for Strategic Planning (RASP) methodology, have evolved to incorporate multi-source urban flooding (Merz et al., 2010). River floods vary spatially; flat valleys experience extensive inundation, whereas narrow valleys face high-velocity flows with severe mechanical impacts (). Extreme rainfall events, such as the 2002 Elbe flood, underscore the need for robust risk analyses to guide infrastructure design and emergency preparedness (; Shah et al., 2018).
Previous flood susceptibility studies have reported varying model performances. In Iran, Rahmati et al. (2016) achieved AUC values of 0.85–0.92 using frequency ratio and logistic regression. In China, Wang et al. (2015) reported an RF AUC of 0.89 for flood hazard assessment. In Southeast Asia, achieved a CNN R2 of 0.91 for fluvial flood prediction. Our ANN performance (R2 = 0.910, RMSE = 0.038) is comparable to or exceeds these benchmarks, despite operating in a data-scarce region, demonstrating the value of integrating multi-source satellite data.
4 Methodology
4.1 Study area
The study area is Kisumu County (shown in Figure 1). It is one of Kenya’s 47 counties and is located along the shores of Lake Victoria (Obiero, 2022). Geographically, it lies between longitudes 34° 30′ E and 35° 20′ E and latitudes 0° 0′ S and 0° 20′ S (OTHOO, 2021). The region experiences a tropical climate, with mean annual rainfall ranging from 1,200 mm to 2,000 mm (). Rainfall occurs in two main seasons, the long rains from March to May and the short rains between September and November, with an average annual rainfall of 450–600 mm during the latter (Odhiambo, 2019).
FIGURE 1
Urbanization, deforestation, and infrastructural development have significantly altered the region’s landscape, leading to unpredictable rainfall patterns, prolonged droughts, reduced arable land, and increased siltation in water bodies (). With a total area of 2,093 km2, the county faces land scarcity, with an average of 0.48 ha per rural resident, a figure that drops to 0.25 ha in some areas (). This population density is a challenge given Kenya’s rapid population growth.
4.2 Data acquisition
4.2.1 Spatial dynamic variables
Flood extent in Kisumu County was assessed by incorporating key spatial dynamic variables that influence flood responses, as summarized in Table 1. These datasets were collected between February 14 and May 14, which corresponds to the rainy season in Kisumu County, for the years 2014–2025. Landsat eight satellite imagery was used for the spectral indices, the normalized difference built-up index (NDBI), and the normalized difference vegetation index (NDVI).
TABLE 1
| Dataset | Formula | Bands | Influence on flood response |
|---|---|---|---|
| NDBI | NDBI = (SWIR1 − NIR)/(SWIR1 + NIR) | B6 = short-wave infrared 1 (SWIR1) band B5 = near infrared (NIR) band | High NDBI values correspond to impervious surfaces that dramatically reduce infiltration capacity. During rainfall, these areas generate rapid surface runoff, increasing peak discharge and shortening the time to flood onset. In urban and peri-urban Kisumu, NDBI serves as a direct proxy for anthropogenically amplified flood hazard. |
| NDVI | NDVI = (NIR − red)/(NIR + red) | B5 = near infrared (NIR) band B4 = red band | NDVI measures live green vegetation density. High NDVI indicates dense vegetation, which intercepts rainfall, increases infiltration through root channels, and promotes evapotranspiration, thereby reducing net runoff volume. Conversely, low NDVI signifies reduced hydrological buffering, leading to higher overland flow. |
| Sentinel1 Flood mask | Flooded_Pixel = (σ°_pre-flood − σ°_flood) > threshold | σ°_pre-flood: median backscatter value (in dB) from the pre-flood period. σ°_flood: median backscatter value (in dB) from the flood period. Threshold: determines the sensitivity of the flood detection | The binary flood mask, created through change detection between pre-flood and flood-phase Sentinel-1 imagery, defines the target variable for the machine learning models. It teaches the algorithms to recognize the specific combination of environmental conditions (soil moisture, topography, and land cover) that lead to inundation. This mask isolates event-based flooding from permanent water bodies, ensuring the model learns transient flood signatures. |
Parameter settings and analysis for the datasets used in flood risk modeling.
4.2.2 Built-up
NDBI was obtained from the Landsat image. Figure 2a shows a crucial dataset that helped in understanding how built environments affect water flow and filtration due to the impervious surface nature of such environments.
FIGURE 2
4.2.3 Soil moisture
Soil moisture is a critical parameter for flood mapping as it influences runoff generation, infiltration capacity, and the likelihood of flood occurrence. The Soil Moisture Active Passive (SMAP) Level-4 soil moisture product shown in Figure 2c was used to derive surface soil moisture conditions in this study. It provides global 3-hourly, 9-km resolution estimates of surface (0–5 cm) and root-zone (0–100 cm) soil moisture by assimilating L-band brightness temperature data into a land surface model. During periods of instrument outages, model simulations ensure data continuity. The dataset is processed and provided in EASE-Grid 2.0.
4.2.4 Vegetation
The normalized difference vegetation index (Figure 2c) shows how the NDVI from the Landsat image was used for vegetation cover, which influences infiltration rates and evapotranspiration. From it, we understand how different vegetation covers affect water flow and filtration. Areas with extensive cultivation are more prone to erosion, which increases surface runoff rather than promoting unrestricted water flow.
4.2.5 Sentinel 1
Radar SAR data are particularly advantageous for flood mapping due to their ability to penetrate cloud cover and operate day and night. Sentinel-1A Radar is an active microwave remote sensing instrument capable of detecting water and providing high-resolution images under all weather, day, and night conditions. Figure 3 shows a map of historical flood extent in Kisumu County during high-rain months between February and May of each year of our study period, 2014-2025. The wide use of SAR systems in flood monitoring lies in the sensitivity of the backscatter signal to open water (Zotou et al., 2020). Flood extents were derived from the Sentinel-1 Ground Range Detected (GRD) collection, available on the Google Earth Engine (GEE) data catalog (COPERNICUS/S1-GRD). C-band SAR data are ideal for detecting surface water due to the specular reflection of radar signals off calm water bodies, which appear as very dark areas (low backscatter) in the imagery. Sentinel-1 is a two-SAR satellite constellation designed to guarantee global coverage with a revisit time of 6 days.
FIGURE 3
The core principle of SAR-based flood mapping is change detection. A pre-flood (dry) reference image is compared to an image acquired during a flood event. A significant drop in backscatter between the two dates indicates potential inundation. A common challenge in flood mapping is the misclassification of permanent water bodies, such as rivers and lakes, as new floods. To address this, the algorithm integrated static water data to increase accuracy. The Joint Research Centre (JRC) Global Surface Water Occurrence dataset JRC/GSW1_3/Global Surface Water was used. This dataset provides the percentage of time each pixel was covered by water over a multi-year period and was used to mask permanent water bodies.
The preliminary flood extent was refined by masking out pixels where the permanent water occurrence was greater than 50%, helping to avoid mislabeling natural water as flooded pixels. This critical step ensures that the final flood extent represents only newly inundated areas, excluding the permanent waters of Lake Victoria and the major rivers within Kisumu County, thereby isolating the actual flood event.
The threshold value in the change detection equation was determined using Otsu’s method for automatic threshold selection. This method analyzes the histogram of backscatter difference values (σ°_pre-flood − σ°_flood) and identifies the optimal threshold that maximizes between-class variance between flooded (low backscatter) and non-flooded (high backscatter) pixels. For our Sentinel-1 GRD collection, the threshold was automatically computed per image scene to account for variations in background backscatter due to soil moisture, vegetation, and incidence angle effects. Pixels with backscatter drop exceeding this threshold were classified as flooded.
4.3 Spatial static variables
Figure 4 shows static variables derived from the Shuttle Radar Topography Mission (SRTM) and DEM, which were incorporated to improve flood risk and prediction mapping. The SRTM V3 dataset, at 30 m resolution, provides accurate elevation data from which slope and stream networks were extracted. Slope was calculated to identify areas of steep gradients where water rapidly drains, reducing flood potential, versus flat lowlands where water tends to accumulate. Stream networks derived from flow direction and accumulation analyses highlight drainage pathways that control flood routing. When combined with the Lake Victoria shoreline shapefile, these variables are critical in delineating flood-prone areas, as regions adjacent to the lake or with gentle slopes exhibit higher inundation risk.
FIGURE 4
4.4 Spatial resolution harmonization
Spatial harmonization is important after clipping the datasets. The raster layers inevitably had slightly different dimensions. A resizing-to-majority algorithm was used to identify the most common dimensions among all available clipped raster datasets. Crucially, the padding dimensions were determined using training data (2014–2022). The “resizing-to-majority” algorithm identified the most common dimensions among training years only. These training-derived dimensions were then applied to pad validation (2023–2025) and forecast (2026–2027) rasters.
The stack shown in Figure 5 is intersected to get the common pixels, where data are available for all raster images. Raster values are then extracted at the training point locations to create ML data (). For each target flood event, a corresponding y-timesteps-day window of meteorological data leading up to the event was extracted. Flexible date handling was implemented to find a continuous block of data for the required period.
FIGURE 5
The final products were generated at 30 m resolution for risk maps (matching SRTM and Landsat) and 10 m resolution for flood extent maps (native Sentinel-1 resolution).
4.4.1 Temporal composition and cloud masking
Gap-filling and temporal interpolation were conducted for years with missing spatial data, specifically NDBI and NDVI. A temporal interpolation scheme was employed for a missing year; the two closest available years, one previous and one next, were identified. A multitemporal image series has the potential to better document change over time as it contains more information from which to estimate missing values, and various approaches have been proposed for filling gaps in multitemporal remote sensing data (Malambo and Heatwole, 2015). It should be noted that some methods focus on predicting missing values and are applied before data analysis, while others are designed to analyze the data and return predictions for missing values as a byproduct. A linear interpolation is always performed based on the temporal distance to fill the gap (). If only past or future data were available, the most recent or immediate future data were replicated, respectively. If there was no data for the recent past or next year, then we used the mean for all available years.
Cloud masking was applied using the CFMASK algorithm available in Google Earth Engine. Pixels with greater than 20% cloud cover were excluded from the annual composite. For partially cloudy scenes, temporal interpolation using adjacent cloud-free years was performed.
4.4.2 Data pre-processing and the patch-based strategy
A critical challenge in flood mapping is the extreme class imbalance; 80.2% non-flood pixels vastly outnumber 19.8% flood pixels, which is due to high flood extents in some areas, especially around Lake Victoria, and very low extents in other areas. Restated, sample distributions are highly skewed because some classes appear rarely compared to others. In the context of flood data, the minority class with a high risk of floods is more interesting from both a learning point and a disaster mitigation perspective. The same problem exists in many other real-world domains, such as medical applications, risk management, detection of fraudulent telephone calls, and biological data analysis (Wang et al., 2024).
Training on full scenes would bias the model toward predicting “no flood.” To solve this, a novel patch-based sampling strategy was implemented, focusing on flooded areas. For each year with a valid flood mask, the coordinates of all flooded pixels were identified. Small, focused patches of 64 × 64 pixels were extracted from all input data layers, centered on these flooded pixels. Area definitions utilize quantitative criteria such as crown density, minimum patch size, or minimum patch width to facilitate a forest/no forest decision in order to construct the sampling frame ().
To ensure the model also learned what non-flood conditions look like, additional patches were sampled from low-flood and medium-flood intensity areas within the annual scenes. This created a final dataset where flood pixels were significantly more represented (∼70–90% in flooded patches), reversing the natural imbalance and forcing the model to learn discriminating features. Landscape relates to all the patches or shapes in a theme. Patches examine individual polygons or a contiguous set of cells (Paudel and Yuan, 2012). The training dataset was artificially expanded by applying random transformations, horizontal/vertical flips, 90-degree rotations, and slight brightness adjustments to the patches, improving model generalization. Based on the annual dataset, the resulting shapes are x (240, 14, 64, 64, 8) representing (samples, timesteps, height, width, features), y (240, 64, 64, 1) representing (samples, height, width, channels).
4.4.3 Ground-truth data description
Radar SAR data are particularly advantageous for flood mapping due to their ability to penetrate cloud cover and operate day and night. The wide use of SAR systems in flood monitoring lies in the sensitivity of the backscatter signal to open water (). Flood extents used as ground truth were derived from the Sentinel-1 Ground Range Detected (GRD) collection, available on the Google Earth Engine (GEE) data catalog (COPERNICUS/S1-GRD).
The core principle of SAR-based ground-truth generation is change detection. A pre-flood (dry) reference image is compared to an image acquired during a flood event. A significant drop in backscatter between the two dates indicates potential inundation. To avoid misclassification of permanent water bodies, the JRC Global Surface Water Occurrence dataset (JRC/GSW1_3/Global Surface Water) was used to mask out pixels where permanent water occurrence exceeded 50%.
For model training and validation, annual flood masks were created for the period 2014–2025, focusing on the rainy season months (February–May). Each mask was manually quality-checked against historical flood records from the Kenya Meteorological Department and local news reports. The final ground-truth dataset consisted of 240 patches of 64 × 64 pixels, with each patch containing approximately 70%–90% flood pixels to address class imbalance. The training period (2014–2022) used 192 patches, while the validation period (2023–2025) used 48 patches.
4.5 Model architecture
4.5.1 Artificial neural network
Artificial neural networks have been considered as an alternative to physically based models due to their simplicity regarding the minimum requirements for collecting detailed data (Wannewitz and Garschagen, 2021). Figure 6 shows the artificial neural network architecture. A multi-layer perceptron architecture was implemented with an input layer (9 neurons), hidden layers (64 and 32 neurons with ReLU activation), and an output layer (1 neuron). In a typical ANN design, the output layer reflects the intended values of the prediction parameter for which the modeling is to be conducted. The output is given from the data during the training phase () using Equation 1:
FIGURE 6
where is the input features, is the weight connecting input I to hidden neuron j, is bias for hidden neuron j, and is the activation function.
4.5.2 Convolutional neural network
A key factor that has made the deep CNNs extremely popular is the ever-increasing computational power of modern computers. Figure 7 shows the 1D convolutional architecture designed to capture spatial patterns in the input data (9 features × 1 channel). Convolutional neural networks, in general, are feed-forward neural networks with alternating convolutional and subsampling layers and are predominantly trained in a supervised manner () using Equation 2.where is the filter weights, is the kernel size, is the input, and is the ReLU activation function.
FIGURE 7
4.5.3 Deep neural network
A multi-layer perceptron (MLP) repeatedly learns the weights and biases between input and output data through mathematical modeling by imitating the neural network of a deep neural network (DNN), an algorithm developed to enable in-depth learning using multiple hidden layers in a conventional MLP having an input layer, a hidden layer, and an output layer (Park and Lee, 2024). Figure 8 shows the deep neural network architecture implemented using TensorFlow/Keras. A deeper neural network architecture was implemented using TensorFlow/Keras with an enhanced input layer capacity of nine features. The hidden network consists of a dense layer (128 neurons, ReLU activation), a dropout layer (30% rate), a dense layer (64 neurons, ReLU activation), a dropout layer (20% rate), a dense layer (32 neurons, ReLU activation), and finally, a single-neuron output layer with linear activation using Equation 3.
FIGURE 8
4.5.4 Random forest
A random forest is an ensemble of classification and regression trees that overcomes overfitting issues of single decision trees while retaining their predictive accuracy. The RF algorithm is known for its reliability in making predictions. Random forest is one of the most established tree-based machine learning methods in flooding. The RF algorithm remains effective even with sizable datasets, demonstrating resilience against overfitting (Wahba et al., 2024). The technique was developed and became a popular tool in many geoscientific fields due to its flexibility and availability in popular software such as R or MATLAB (Schoppa et al., 2020).
A random forest model was implemented as an ensemble learning method that constructs multiple decision trees. For our case, the ensemble was configured with 100 trees, five levels per tree, a minimum of two samples required to split, and a minimum of one sample at a leaf. Node splitting was evaluated using the mean squared error criterion, and a random state of 42 was set to ensure reproducibility during training. The ensemble outputs the mean prediction of the individual trees. The architecture was designed with these specific parameters to balance complexity and prevent overfitting.
The RF algorithm operates on the principle of bootstrap aggregating (bagging), where multiple decision trees are trained on different subsets of the data. Each tree is grown using a random subset of features, and the final prediction is obtained by averaging the predictions of all trees. This approach reduces variance while maintaining low bias, making it particularly effective for tabular data with mixed feature types.
For a random forest with T trees, the prediction for input x is given by using Equation 4
where ft(x) is the prediction of the tth decision tree.
4.5.5 Ensemble model
Ensemble prediction allows for representation of the forecast uncertainty by producing an ensemble of possible forecast outcomes, each of which is equally probable (Wu et al., 2020b). The move toward ensemble prediction systems (EPS) in flood forecasting represents the state of the art in forecasting science, following on the success of the use of ensemble design for weather forecasting and paralleling the move toward ensemble forecasting in other related disciplines such as climate change predictions (). Each of these approaches exhibits certain weaknesses in terms of flood risk mapping that can be improved through an ensemble modeling approach. Ensemble modeling is a process of combining the predictions of single models into an integrated model to increase prediction accuracy (). There are three preferred ensemble methods: boosting, bagging, and stacking. Boosting methods assemble homogeneous types of multiple models that learn to fix their earlier prediction errors. Bagging generates multiple models from several subsamples of the training dataset (Prasad et al., 2022).
A model averaging ensemble was constructed by combining results from all four individual models using this approach using Equation 5 :
Equal weighting (1/4 for each model) was selected as a conservative baseline to avoid overfitting to validation data, consistent with recommendations for ensemble flood forecasting where model performance is comparable, and overfitting risks outweigh marginal gains from optimized weights (Wu et al., 2020a).
4.5.6 Model training
The models utilized nine forecasting features across three primary categories: a temporal feature (years from 2014 to 2025), spatial features (NDBI, NDVI, soil moisture, slope, and stream proximity), and meteorological features (temperature, rainfall, and wind speed). The dataset was partitioned using a time-based split, with 80% allocated for training and 20% held out for validation.
The random forest model was trained using bootstrap sampling with replacement, considering the square root of the total number of features at each split. No additional hyperparameter tuning was performed beyond this initial configuration.
Neural networks were initialized with a normal distribution for weights in layers with ReLU activations and trained using mini-batch gradient descent. Training configurations included batch sizes of 8–16 samples depending on model complexity, a maximum of 200 epochs, and early stopping with a patience of 20 epochs to monitor validation loss and restore the best weights, alongside a reduction in learning rate on plateau. All models used a fixed random seed of 42 for full reproducibility.
To mitigate overfitting, multiple regularization techniques were employed: Dropout was applied in DNN and CNN architectures, randomly disabling 20%–30% of neurons during training, and L2 regularization (with α = 0.001) was used for the ANN, which adds a penalty term to the loss function proportional to the sum of squared weights using Equation 6:
The scaling of forecast data was a critical step to ensure consistency with the model’s training paradigm, governed by Equation 7:
where Z is the scaled value, X is the original value, μ is the mean, and σ is the standard deviation.
To prevent data leakage and preserve the chronological integrity of the forecasting framework, feature scaling was performed using a strict temporally separated procedure. No information from future years was used in estimating the normalization parameters. Specifically, a StandardScaler was fitted exclusively on the training dataset, from which the feature-wise mean and standard deviation were computed. These statistics were then retained and applied unchanged to all subsequent datasets. The validation period was transformed using the training-derived parameters according to according to Equation 8
All models output a continuous probability value between 0 and 1 for each pixel, representing the likelihood of flood inundation. This probability is then classified into five risk categories for mapping purposes using Jenks natural breaks classification.
5 Results and discussion
5.1 Spatial distribution of flood risk categories
A spatial analysis of flood risk in Kisumu County reveals varying degrees of risk across the region. The assessment classified the study area into five risk categories using Jenks natural breaks classification: very high (>0.75), high (0.60–0.75), moderate (0.40–0.60), low (0.20–0.40), and very low (<0.20), based on environmental, hydrological, and meteorological factors, as shown in Figure 9.
FIGURE 9
Flood risk mapping (Figure 10) of Kisumu County was undertaken using ANN, CNN, DNN, and RF. The results reveal considerable variation in the allocation of risk categories across the models. The ANN model predicted that 2055.83 km2 could be classified as very low risk, with only 14.57 km2 identified as very high risk, and small areas assigned to moderate and high classes. Similarly, the RF model mapped 1905.17 km2 as very low risk, 165.69 km2 as high risk, and 6.82 km2 as very high risk. The CNN model showed a strong bias, classifying 1900.58 km2 as very low risk and only 6.49 km2 as very high, with negligible coverage in the other categories. In contrast, the DNN model produced a more distributed classification, identifying 811.0 km2 as moderate risk, 467.0 km2 as high risk, and 72.01 km2 as very high risk. These results suggest that while ANN, RF, and CNN favor conservative classifications, the DNN model demonstrates greater sensitivity in detecting flood-prone areas.
FIGURE 10
In Figure 11, the very high-risk category is extremely limited, covering only 19.2 km2, which constitutes 0.9% of the county along the shores of Lake Victoria and the areas around Katito, indicating isolated hotspots highly prone to flooding. The high-risk areas span 220.3 km2, 10.5% of total area coverage, with a very close focus being on Kisumu City, Ahero, and Katito, highlighting regions that require priority in flood mitigation planning. The moderate-risk category dominates the flood-prone zones, encompassing 618.3 km2, 29.6% of total area coverage, suggesting that nearly one-third of Kisumu County is moderately vulnerable to flood events. These observations are closely linked to stream proximity, lake proximity, built environments, and flat terrains.
FIGURE 11
On the other hand, the low-risk areas cover 795.1 km2 (38.1%), representing the largest portion of the county with minimal flood exposure, likely including elevated terrain and regions with effective drainage. Very-low-risk zones account for 438.7 km2 (21.0%), areas virtually unaffected by flooding under typical climatic conditions, mostly areas around Maseno and Riat in Kisumu County, with Figure 10 showing risk locations.
Overall, the flood risk distribution indicates that while Kisumu County has significant areas at moderate to high risk, approximately 40.5% combined, most of the region remains at low or very low risk. These insights are essential for local authorities to prioritize flood prevention measures, land-use planning, and emergency response strategies.
This spatial pattern aligns strongly with established knowledge of flood drivers. The high risk along the Lake Victoria shoreline and river systems underscores the paramount importance of proximity to water bodies as a primary flood driver, a factor consistently identified in flood susceptibility studies globally (Rahmati et al., 2016; Wang et al., 2015). The concentration of high risk in urban centers such as Kisumu City is a direct consequence of anthropogenic factors. The NDBI values in these areas indicate significant impervious surfaces, which reduce infiltration and accelerate surface runoff, a phenomenon well-documented in urban flood literature (Pradhan, 2010; ). The identification of Ahero and other flat, low-lying areas as high-risk zones highlights the controlling role of topography and slope, where gentle gradients cause water to accumulate rather than drain away (Nigusse and Adhanom, 2019; Yalcin and Akyurek, 2020).
5.2 Validation results
The validation results demonstrated exceptional predictive performance across most models, as shown in Figure 12. The artificial neural network approach achieved the highest individual performance, with an R2 value of 0.910 and RMSE of 0.038, indicating that the model explains 91% of the variance in flood risk with minimal error. The CNN and DNN models also showed strong performance with R2 values of 0.880 and 0.890, respectively. The ensemble model, which combined predictions from all individual models, achieved an R2 of 0.800 and RMSE of 0.040, representing a robust and conservative prediction approach.
FIGURE 12
In Table 2, the RF model’s relatively lower performance (R2 = 0.520) highlights a key limitation of tree-based models for this specific application. RF operates by creating axis-aligned splits in the feature space, forming a series of decision boundaries. This approach struggles with smoothly varying continuous functions and complex interaction effects that are characteristic of hydrological processes. For instance, the relationship between soil moisture and flood probability is not a simple threshold but a continuous function that interacts with rainfall density and NDBI. Neural networks’ ability to model smooth, interactive decision boundaries through weighted sums and non-linear activations gives them a distinct advantage. The exceptional performance of these models, particularly the ANN with an R2 value of 0.910, provides high confidence in the forecasting capabilities of the implemented framework for flood risk assessment in the study region.
TABLE 2
| Model | ANN | CNN | DNN | Random forest | Ensemble model |
|---|---|---|---|---|---|
| R2 | 0.910 | 0.880 | 0.890 | 0.520 | 0.800 |
| RMSE | 0.038 | 0.044 | 0.042 | 0.089 | 0.040 |
Model performance summary.
Excluding random forest from the ensemble consistently improved forecasting performance. The ensemble without RF (ANN + CNN + DNN) achieved an R2 value of 0.916 and reduced RMSE from 0.040 to 0.035, representing a 22.0% improvement in prediction error. This indicates that random forest, with its substantially lower individual performance (R2 = 0.520), introduced systematic bias that degraded the ensemble’s collective predictive capability.
5.3 Spatial accuracy assessment
Beyond explaining variance in flood extent (R2), the spatial accuracy of the ensemble model’s flood forecasting was evaluated using a confusion matrix comparing forecasted flood pixels against observed flood extent data for the validation period. The results are shown in Table 3:
TABLE 3
| Metric | Value |
|---|---|
| Overall accuracy | 0.80 |
| Kappa coefficient | 0.78 |
| Precision | 0.83 |
| Recall | 0.80 |
| F1-score | 0.81 |
| ROC-AUC | 0.94 |
Spatial classification performance metrics.
These metrics confirm that the model not only captures temporal flood variability but also accurately identifies the spatial locations of inundation, particularly in high-risk zones along the Lake Victoria shoreline and in Kisumu City.
Ground-truth validation was performed using historical flood extent derived from Sentinel-1 SAR imagery processed independently of the model training pipeline. Comparison against documented flood events from Kenya Meteorological Department records and local news reports showed 85% agreement, with the model successfully capturing the spatial extent of major flood events, including the 2014, 2018, and 2020 inundations in Kisumu County.
5.4 Feature importance
The analysis of feature importance shown in Figure 13 for flood risk mapping provided critical insights into the dominant drivers in Kisumu County. Slope emerged as the most influential factor (20%), confirming that topography is a primary control on flood water accumulation and flow velocity. This aligns with numerous studies that identify slope as a fundamental variable in flood susceptibility index models (; Singha et al., 2024).
FIGURE 13
Anthropogenic and hydrological factors: NDBI (15%), soil moisture (15%), stream proximity (15%), and lake proximity (15%) were equally critical. The high importance of NDBI reinforces the role of urbanization in exacerbating flood risk by increasing surface runoff (Pradhan, 2010). The significance of soil moisture highlights the role of antecedent conditions, where saturated soils significantly increase flood potential by reducing infiltration capacity, a key concept in hydrology (Wasko et al., 2021; ). Proximity to streams and the lake directly measures exposure to the flood hazard source, a classic and persistent factor in flood risk assessment (Wang et al., 2015).
Notably, rainfall contributed 10%, signifying its role as a direct trigger, but its effect is heavily moderated by these static and dynamic landscape characteristics. This finding is crucial as it moves beyond a simplistic “heavy rain causes floods” narrative to a more nuanced understanding that the landscape’s pre-condition and structure largely determine the impact of a rainfall event ().
5.5 Uncertainty quantification
Reliable flood risk assessment requires not only accurate forecasting but also robust characterization of forecast uncertainty. Decision makers need to know not only what will happen, but also how confident they can be in that prediction. Following best practices in ensemble flood forecasting (Wu et al., 2020b; ), we quantified uncertainty using three complementary approaches: bootstrap resampling for model variance, forecast interval estimation, and decomposition of aleatoric versus epistemic uncertainty.
5.6 Bootstrap variance estimation
To quantify the stability and variance of our model forecasting, we employed bootstrap resampling with replacement. For each model (ANN, CNN, DNN, and the ensemble), we generated 10 bootstrap samples from the training dataset. Each bootstrap sample contained the same number of observations as the original training set (n = 240 patches) but with random replacement, creating slightly different training distributions. We retrained each model on all 10 bootstrap samples and recorded the predictions for the validation set. The results are presented in Table 4.
TABLE 4
| Model | Mean R2 | R2 Std. Dev. | Mean RMSE | RMSE Std. Dev. | 95% confidence interval (R2) |
|---|---|---|---|---|---|
| ANN | 0.91 | 0.008 | 0.038 | 0.002 | 0.892–0.924 |
| CNN | 0.88 | 0.012 | 0.044 | 0.003 | 0.852–0.900 |
| DNN | 0.89 | 0.010 | 0.042 | 0.003 | 0.867–0.907 |
| Ensemble | 0.916 | 0.006 | 0.036 | 0.002 | 0.904–0.928 |
Bootstrap variance estimates.
The low standard deviations across bootstrap runs (R2 σ < 0.012 for all models) indicate that our models are stable and not overly sensitive to sampling variability. The ensemble model shows the smallest variance (R2 σ = 0.006), demonstrating that combining multiple architectures reduces forecasting instability, a key advantage of ensemble methods (Prasad et al., 2022).
5.7 Prediction intervals
Beyond point estimates, we constructed prediction intervals to communicate the range within which future flood extents are likely to fall. For the ensemble model, forecast intervals were derived from the distribution of individual model predictions at each pixel. Following standard practice (
Tang et al., 2023), we report:
±1σ interval (68% confidence): The range containing approximately 68% of likely outcomes
±2σ interval (95% confidence): The range containing approximately 95% of likely outcomes
Figure 14 presents the ensemble forecast for 2026–2027 with confidence bands. The widening of confidence bands over the forecast horizon reflects increasing uncertainty with longer lead times, a phenomenon well-documented in hydrological forecasting (Vilanculos, 2015). For 2026, the ±1σ interval ranges from 0.35% to 0.45% flood extent, while the ±2σ interval extends from 0.31% to 0.49%. By 2027, uncertainty expands slightly (95% CI: 0.33%–0.52%), reflecting cumulative errors in meteorological inputs and model approximations.
FIGURE 14
5.8 Aleatoric vs. epistemic uncertainty decomposition
Following established frameworks for uncertainty quantification in environmental modeling (Yariyan et al., 2020), we decomposed total prediction uncertainty into two components, aleatoric uncertainty and epistemic uncertainty.
5.8.1 Aleatoric uncertainty (irreducible)
This uncertainty stems from inherent randomness in the flood-generating process, stochastic rainfall variability, measurement errors in satellite data, and natural fluctuations in soil moisture. Aleatoric uncertainty cannot be reduced by collecting more data or improving the model; it represents the fundamental unpredictability of the system. We estimated aleatoric uncertainty by measuring the residual variance between model predictions and observed flood extents after accounting for all available features.
5.8.2 Epistemic uncertainty (reducible)
This uncertainty arises from limitations in our knowledge, model structure choices, parameter estimation errors, and incomplete representation of physical processes. Epistemic uncertainty can be reduced through better data, improved model architectures, and more rigorous validation. We quantified epistemic uncertainty as the variance between different model predictions (ANN, CNN, DNN) for the same input, following the “model disagreement” approach (Wu et al., 2020a). The results are presented in Table 5.
TABLE 5
| Uncertainty type | Contribution to total variance | Description |
|---|---|---|
| Aleatoric | 62% | Inherent randomness in rainfall, soil moisture dynamics |
| Epistemic | 38% | Model structure differences, parameter uncertainty |
Results of decomposition.
5.9 Monte Carlo dropout for neural network uncertainty
For neural network models (ANN, CNN, and DNN), we implemented Monte Carlo (MC) dropout as an additional uncertainty quantification technique (; ). MC dropout exploits dropout regularization, typically used only during training, by keeping dropout active during inference. By performing multiple forward passes (N = 50) with different random dropout masks, we generate a distribution of predictions for each input.
5.10 Methods
For each test sample, we performed 50 stochastic forward passes through each neural network with dropout enabled (30% dropout rate for DNN, 20% for ANN). The mean forecast served as the point estimate, while the standard deviation across passes quantified model uncertainty. This approach approximates Bayesian inference in deep neural networks at low computational cost. The results are presented in
Table 6.
Transition zones between flooded and non-flooded pixels (land–water boundaries).
Complex urban areas where drainage patterns are heterogeneous.
Flat, low-lying regions where small topographic variations determine flood extent.
TABLE 6
| Model | Mean prediction uncertainty (σ) | High-uncertainty regions |
|---|---|---|
| ANN | 0.008 | Transition zones (land-water boundaries) |
| CNN | 0.011 | Urban areas with complex drainage |
| DNN | 0.009 | Flat terrains near Lake Victoria |
Key findings from MC dropout.
MC dropout revealed that prediction uncertainty is not uniformly distributed across the study area. The highest uncertainty occurs in
These findings align with expectations from hydraulic theory (Liu et al., 2025; Peramuna et al., 2025) and provide actionable guidance for targeted data collection: improving topographic resolution in transition zones would yield the greatest uncertainty reduction.
5.11 Recommendations
Operationalizing the findings is crucial for saving lives and protecting livelihoods in the short and medium term.
5.11.1 Operationalize the predictive early warning system
The high concentration of people and assets at risk necessitates a robust early warning system. The developed ensemble model must be transitioned into an operational flood early warning system (FEWS) by creating a real-time data pipeline that integrates live meteorological and satellite data.
5.11.2 Enhance targeted community-based early warning
Forecasts and warnings can be disseminated through multiple channels (SMS, radio, and community loudspeakers), tailored to reach the vulnerable populations identified in the vulnerability map. Regular community drills could be conducted to ensure understanding and prompt action.
5.11.3 Invest in nature-based solutions for risk reduction
To mitigate flood peaks and increase infiltration, strategic investments should be made in nature-based solutions (NBSs). This includes reforestation in upstream catchment areas, the restoration and conservation of wetlands and riparian zones to act as natural sponges, and the creation of green spaces within urban areas to reduce impervious surface cover.
5.11.4 Incorporate flood risk maps into mandatory land-use planning
The generated flood risk maps should be legally mandated tools for all physical planning and development control processes. Spatial zoning regulations must prohibit new critical infrastructure (schools, hospitals) and high-density residential developments in the identified very high and high-risk zones.
5.12 Limitations of our study
1. While Sentinel-1 SAR data enables all-weather flood mapping, its performance is challenged in densely vegetated areas. Vegetation causes signal attenuation and volume scattering, which reduces the contrast between flooded and non-flooded pixels, potentially leading to false negatives (underestimation of flood extent). Our integration of NDVI partially mitigates this limitation by providing complementary information on vegetation density. However, in regions with dense riparian vegetation or forested wetlands, users should interpret flood extents with caution. Future work should consider integrating optical imagery, such as Sentinel-2 or polarimetric SAR decomposition techniques, to improve flood detection in vegetated floodplains.
2. Model generalizability and temporal scope. The model was trained on a specific 12-year historical period (2014–2025) from Kisumu County. Its direct application to a different region or under drastically different future climate regimes would require re-training and validation. To enhance generalizability, the model could be tested and fine-tuned in other Lake Victoria counties (Homa Bay, Siaya, and Busia). Incorporating future climate projection data (CMIP6) would also allow for “what-if” scenario analysis under different climate change pathways.
6 Conclusion
This study successfully developed and validated a high-resolution flood risk assessment and forecasting framework for Kisumu County, Kenya, by integrating multi-source geospatial data (Sentinel-1 SAR, Landsat 8, SRTM DEM, and SMAP soil moisture) with advanced machine learning architectures (ANN, CNN, DNN, random forest, and ensemble methods). The research addressed critical gaps in data-scarce regions, demonstrating that satellite-derived environmental variables can reliably model complex flood dynamics without reliance on extensive in situ hydrological networks.
The spatial analysis produced a high-resolution flood risk map (Figure 9) classifying Kisumu County into five distinct risk categories based on Jenks natural breaks thresholds. While 59.1% of the county falls within low or very low risk zones, a substantial 40.9% (857.8 km2) faces moderate to very high flood risk. Critically, very high-risk zones, although comprising only 0.9% (19.2 km2) of the total area, were concentrated along the shores of Lake Victoria and near Katito, aligning precisely with historical flood records. High-risk areas (10.5%, 220.3 km2) are notably focused around Kisumu City and Ahero, exposing densely populated urban and peri-urban settlements to recurrent inundation hazards.
Among the four machine learning models evaluated, the artificial neural network achieved the highest predictive accuracy (R2 = 0.910, RMSE = 0.038), demonstrating exceptional capability to capture the complex, non-linear relationships between environmental drivers and flood occurrence. Excluding random forest from the ensemble (due to its systematic underprediction bias, R2 = 0.520) improved overall performance, with the ANN–CNN–DNN ensemble achieving an R2 of 0.916 (RMSE = 0.035), a 22% reduction in prediction error compared to the full ensemble. Feature importance analysis identified slope (20%), NDBI (15%), soil moisture (15%), and stream proximity (15%) as the dominant flood drivers, confirming that topography and anthropogenic land use moderate flood risk more strongly than rainfall intensity alone.
Bootstrap resampling (10 runs) revealed low model variance (ensemble R2 σ = 0.006), confirming prediction stability. Forecasts for 2026–2027 predict a slight but consistent increase in flood extent from 0.397% to 0.441% of the county area, with 95% confidence intervals widening from ±0.09% to ±0.10% over the forecast horizon. Aleatoric uncertainty (62% of total variance) dominated epistemic uncertainty (38%), indicating that while inherent rainfall stochasticity fundamentally limits predictability, targeted improvements in topographic data and real-time rainfall networks could meaningfully reduce forecast uncertainty.
While socioeconomic vulnerability analysis was beyond the scope of this manuscript, the spatial risk map provides a foundation for such assessments in future work. Without targeted intervention, the projected increase in flood extent, coupled with ongoing urbanization and climate-induced rainfall intensification in the Lake Victoria Basin, will likely exacerbate disaster impacts. The operationalization of this forecasting framework as a real-time early warning system, combined with mandatory land-use planning informed by the risk maps, represents the most urgent pathway for reducing flood vulnerability in Kisumu County. This study provides a replicable methodology for data-scarce regions globally, demonstrating that machine learning, when coupled with open-access satellite data, can deliver actionable flood intelligence for climate adaptation planning.
Statements
Data availability statement
The original contributions presented in the study are included in the article/supplementary material; further inquiries can be directed to the corresponding author.
Author contributions
HA: Conceptualization, Data curation, Formal Analysis, Funding acquisition, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review and editing. MG: Supervision, Writing – review and editing. AO: Writing – review and editing.
Funding
The author(s) declared that financial support was not received for this work and/or its publication.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
References
1
AhmadM.AhmedZ.YangX.HussainN.SinhaA. (2022). Financial development and environmental degradation: do human capital and institutional quality make a difference?Gondwana Res.105, 299–310. 10.1016/j.gr.2021.09.012
2
Al-NoutiA. F.FuM.BokdeN. D. (2024). Reservoir operation based machine learning models: comprehensive review for limitations, research gap, and possible future research direction. Knowledge-Based Eng. Sci.5 (2), 75–139. 10.51526/kbes.2024.5.2.75-139
3
AlfieriL.ThielenJ.PappenbergerF. (2012). Ensemble hydro-meteorological simulation for flash flood early detection in southern Switzerland. J. Hydrol.424, 143–153. 10.1016/j.jhydrol.2011.12.038
4
AmatebelleC. E.OwolabiS. T.OgundejiA. A.OkolieC. C. (2025). A systematic analysis of remote sensing and geographic information system applications for flood disaster risk management. J. Spat. Sci.70, 1–27.
5
AvandM.MoradiH. R.Ramazanzadeh LasboyeeM. (2021). Spatial prediction of future flood risk: an approach to the effects of climate change. Geosciences11 (1), 25. 10.3390/geosciences11010025
6
BaanP. J.KlijnF. (2004). Flood risk perception and implications for flood risk management in the Netherlands. Int. J. River Basin Manag.2 (2), 113–122. 10.1080/15715124.2004.9635226
7
BakerV. R. (1994). “Geomorphological understanding of floods,” in Geomorphology and Natural Hazards (Amsterdam: Elsevier), 139–156.
8
BankoffS. G.LeeS. C. (1986). A critical review of the flooding literature. Multiph. Sci. Technol.2 (1-4), 95–180. 10.1615/multscientechn.v2.i1-4.20
9
BashaE. A.RavelaS.RusD. (2008). “Model-based monitoring for early warning flood detection,” in Proceedings of the 6th ACM Conference on Embedded Network Sensor Systems. dl.acm. 1–10.
10
BeckerA.GrünewaldU. (2003). Flood Risk in Central Europe. Washington, DC: American Association for the Advancement of Science, 1099.
11
BlomC.VoesenekL. (1996). Flooding: the survival strategies of plants. Trends Ecol. Evol.11 (7), 290–295. 10.1016/0169-5347(96)10034-3
12
BlöschlG.Ardoin‐BardinS.BonellM.DorningerM.GoodrichD.GutknechtD.et al (2007). At what scales do climate variability and land cover change impact on flooding and low flows?Hydrol. Processes21 (9), 1241–1247. 10.1002/hyp.6669
13
BossaA. Y.HounkpèJ.YiraY.SerpantiéG.LidonB.FusillierJ. L.et al (2020). Managing new risks of and opportunities for the agricultural development of West-African floodplains: hydroclimatic conditions and implications for rice production. Climate8 (1), 11. 10.3390/cli8010011
14
BrunnerM. I.SlaterL.TallaksenL. M.ClarkM. (2021). Challenges in modeling and predicting floods and droughts: a review. Wiley Interdiscip. Rev. Water8 (3), e1520. 10.1002/wat2.1520
15
CabreraJ. S.LeeH. S. (2020). Flood risk assessment for Davao Oriental in the Philippines using geographic information system‐based multi‐criteria analysis and the maximum entropy model. J. Flood Risk Manag.13 (2), e12607. 10.1111/jfr3.12607
16
ChenJ.LiY.ShuL.FangS.YaoJ.CaoS.et al (2023). The influence of the 2022 extreme drought on groundwater hydrodynamics in the floodplain wetland of Poyang Lake using a modeling assessment. J. Hydrol.626, 130194. 10.1016/j.jhydrol.2023.130194
17
ChoubinB.MoradiE.GolshanM.AdamowskiJ.Sajedi-HosseiniF.MosaviA. (2019). An ensemble prediction of flood susceptibility using multivariate discriminant analysis, classification and regression trees, and support vector machines. Sci. Total Environ.651, 2087–2096. 10.1016/j.scitotenv.2018.10.064
18
ClokeH. L.PappenbergerF. (2009). Ensemble flood forecasting: a review. J. Hydrol.375 (3-4), 613–626. 10.1016/j.jhydrol.2009.06.005
19
DandapatK.PandaG. K. (2017). Flood vulnerability analysis and risk assessment using analytical hierarchy process. Model. Earth Syst. Environ.3 (4), 1627–1646. 10.1007/s40808-017-0388-7
20
De MoelH.JongmanB.KreibichH.MerzB.Penning-RowsellE.WardP. J. (2015). Flood risk assessments at different spatial scales. Mitig. Adapt. Strategies Glob. Change20 (6), 865–890. 10.1007/s11027-015-9654-z
21
DoubenK. J. (2006). Characteristics of river floods and flooding: a global overview, 1985–2003. Irrigat. Drainage J. Int. Comm. Irrigat. Drainage55 (S1), S9–S21. 10.1002/ird.239
22
GayaC. O. (2020). Application of GIS and Remote Sensing in Flood Management in the Lake Victoria Basin. Juja: JKUAT-COETEC.
23
GerberF.de JongR.SchaepmanM. E.Schaepman-StrubG.FurrerR. (2018). Predicting missing values in spatio-temporal remote sensing data. IEEE Trans. Geoscience Remote Sens.56 (5), 2841–2853. 10.1109/tgrs.2017.2785240
24
HajjiS.KrimissaS.AbdelrahmanK.BoudharA.ElalouiA.IsmailiM.et al (2025). Enhancing flood prediction through remote sensing, machine learning, and Google Earth Engine. Front. Water7, 1514047. 10.3389/frwa.2025.1514047
25
Hossein MojaddadiB. P.NampakH.AhmadN.GhazaliA. H. B. (2017). Ensemble machine-learning-based geospatial approach for flood risk assessment using multi-sensor remote-sensing data and GIS. Geomatics, Nat. Hazards Risk8, 1–20. 10.1080/19475705.2017.1285819
26
Jay Krishna ThakurS. K. S. V. S. E. (2017). Integrating Remote Sensing, Geographic Information Systems and Global Positioning System Techniques with Hydrological Modeling, 7. Berlin: Springer.
27
JonkmanN. T.KalbitzK.BergsmaH.JansenB. (2023). Site history’s role in urban agriculture: a case Study in Kisumu, Kenya, and Ouagadougou, Burkina Faso. Land12 (11), 2056. 10.3390/land12112056
28
KabirS.PatidarS.XiaX.LiangQ.NealJ.PenderG. (2020). A deep convolutional neural network model for rapid prediction of fluvial flood inundation. J. Hydrol.590, 125481. 10.1016/j.jhydrol.2020.125481
29
KarimanziraD.WeisJ.WunschA.RitzauL.LieschT.OhmerM. (2023). Application of machine learning and deep neural networks for spatial prediction of groundwater nitrate concentration to improve land use management practices. Front. Water5, 1193142. 10.3389/frwa.2023.1193142
30
KelleherC.McPhillipsL. (2020). Exploring the application of topographic indices in urban areas as indicators of pluvial flooding locations. Hydrol. Processes34 (3), 780–794. 10.1002/hyp.13628
31
KilybayA.GhoshB.Chacko ThomasN. (2017). A review on the progress of ion‐engineered water flooding. J. Petroleum Eng.2017 (1), 7171957–7171959. 10.1155/2017/7171957
32
KimG.BarrosA. P. (2001). Quantitative flood forecasting using multisensor data and neural networks. J. Hydrol.246 (1-4), 45–62. 10.1016/s0022-1694(01)00353-5
33
KlijnF.KreibichH.de MoelH.Penning-RowsellE. (2015). Adaptive flood risk management planning based on a comprehensive flood risk conceptualisation. Mitig. Adaptation Strategies Global Change20 (6), 845–864. 10.1007/s11027-015-9638-z
34
KöhlM.MagnussenS.MarchettiM. (2006). Sampling Methods, Remote Sensing and GIS Multiresource Forest Inventory. Berlin: Springer.
35
KönigC.-D. (1998). Politisches Handeln Der Städtischen Armen in Kenya, 29. Münster: LIT Verlag.
36
KronW. (2005). Flood risk= hazard• values• vulnerability. Water Internat.30 (1), 58–68. 10.1080/02508060508691837
37
KüblerS.OwengaP.RucinS.C.P. KingG. (2014). “Edaphics, active tectonics and animal movements in the Kenyan Rift-implications for early human evolution and dispersal,” in EGU General Assembly Conference Abstracts (EGU General Assembly).
38
KumarS. S. K. J. (2020). Challenges and recent developments in flood forecasting in India. Roorkee Water Conclave, 163.
39
KundzewiczZ. W.SzwedM.PińskwarI. (2019). Climate variability and floods—A global review. Water11 (7), 1399. 10.3390/w11071399
40
LajiA.AyongaJ. N. (2024). Mainstreaming resilience to flood risk among households in informal settlements in Kisumu City, Kenya. Urban Resil. Sustain.2 (4), 326–347. 10.3934/urs.2024017
41
LastraJ.FernándezE.Díez-HerreroA.MarquínezJ. (2008). Flood hazard delineation combining geomorphological and hydrological methods: an example in the Northern Iberian Peninsula. Nat. Hazards45 (2), 277–293. 10.1007/s11069-007-9164-8
42
LawalZ. K.YassinH.ZakariR. Y. (2021). “Flood prediction using machine learning models: a case study of Kebbi state Nigeria,” in 2021 IEEE Asia-Pacific Conference on Computer Science and Data Engineering (CSDE) (New York, NY: IEEE).
43
LiY.GrimaldiS.WalkerJ. P.PauwelsV. R. N. (2016). Application of remote sensing data to constrain operational rainfall-driven flood forecasting: a review. Remote Sens.8 (6), 456. 10.3390/rs8060456
44
LiZ.WangC.EmrichC. T.GuoD. (2018). A novel approach to leveraging social media for rapid flood mapping: a case study of the 2015 South Carolina floods. Cartogr. Geogr. Inf. Sci.45 (2), 97–110. 10.1080/15230406.2016.1271356
45
LiuW.ZhangX.FengQ.EngelB. A. (2025). City-scale integrated flood risk prediction under future climate change and urbanization based on the shared socioeconomic pathways (SSP) scenarios. J. Hydrol.655, 132971. 10.1016/j.jhydrol.2025.132971
46
MalamboL.HeatwoleC. D. (2015). A multitemporal profile-based interpolation method for gap filling nonstationary data. IEEE Trans. Geoscience Remote Sens.54 (1), 252–261. 10.1109/tgrs.2015.2453955
47
MaseseA.NeyoleE.OmbachiN. (2016). Loss and damage from flooding in lower Nyando Basin, Kisumu County, Kenya. Int. J. Soc. Sci. Humanit. Res.4 (3), 9–22.
48
MasimbeT. (2018). “Impact of climate extremes on water quality and supply in urban informal settlements,” in Kenya: A Case Study of Kisumu City (Nairobi: University of Nairobi).
49
MasonD. C.SpeckR.DevereuxB.SchumannG. P.NealJ.BatesP. (2009). Flood detection in urban areas using TerraSAR-X. IEEE Transactions Geoscience Remote Sensing48 (2), 882–894. 10.1109/tgrs.2009.2029236
50
Md. Shahinoor RahmanL. D. (2017). The state of the art of spaceborne remote sensing in flood management. Nat. Hazards59, 1–15. 10.1007/s11069-017-2997-5
51
MehmoodH.RasmyM. (2020). Challenges and Technical Advances in Flood Early Warning Systems (Fewss). Flood Impact Mitigation and Resilience Enhancement, London: Springer .19.
52
MerzB.ThiekenA.GochtM. (2007). “Flood risk mapping at the local scale: concepts and challenges,” in Flood Risk Management in Europe: Innovation in Policy and Practice (Berlin: Springer), 231–251.
53
MerzB.HallJ.DisseM.SchumannA. (2010). Fluvial flood risk management in a changing world. Nat. Hazards Earth Syst. Sci.10 (3), 509–527. 10.5194/nhess-10-509-2010
54
MoftakhariH. R.AghaKouchakA.SandersB. F.AllaireM.MatthewR. A. (2018). What is nuisance flooding? Defining and monitoring an emerging challenge. Water Resour. Res.54 (7), 4218–4227. 10.1029/2018wr022828
55
MoghimS.GharehtoraghM. A.SafaieA. (2023). Performance of the flood models in different topographies. J. Hydrol.620, 129446. 10.1016/j.jhydrol.2023.129446
56
MunawarH. S.HammadA. W.WallerS. T. (2022). Remote sensing methods for flood prediction: a review. Sensors22 (3), 960. 10.3390/s22030960
57
MunyiJ.-M. M. (2024). The Influence of Urban Morphology on Flood Susceptibility in Slums in A Data Scarce Environment Using Machine Learning. Enschede: University of Twente, 30.
58
NevoS.MorinE.Gerzi RosenthalA.MetzgerA.BarshaiC.WeitznerD.et al (2022). Flood forecasting with machine learning models in an operational framework. Hydrology Earth Syst. Sci.26 (15), 4013–4032. 10.5194/hess-26-4013-2022
59
NguyenD. T.ChenS.-T. (2020). Real-time probabilistic flood forecasting using multiple machine learning methods. Water12 (3), 787. 10.3390/w12030787
60
NguyenM. T.SebesvariZ.SouvignetM.BachoferF.BraunA.GarschagenM.et al (2021). Understanding and assessing flood risk in Vietnam: current status, persisting gaps, and future directions. J. Flood Risk Manag.14 (2), e12689. 10.1111/jfr3.12689
61
NigusseA. G.AdhanomO. G. (2019). Flood hazard and flood risk vulnerability mapping using geo-spatial and MCDA around Adigrat, Tigray region, Northern Ethiopia. Momona Ethiop. J. Sci.11 (1), 90–107. 10.4314/mejs.v11i1.6
62
NjeruL. W. (2025). The Coverage, Framing and Audience Interaction with Climate Change Reporting: A Case Study of Nation. Africa’s Stories. Nairobi: Aga Khan University.
63
NoymaneeJ.TheeramunkongT. (2019). Flood forecasting with machine learning technique on hydrological modeling. Procedia Comput. Sci.156, 377–386. 10.1016/j.procs.2019.08.214
64
ObieroS. O. (2022). Application of Gis and Remote Sensing Methods in Land Use and Land Cover Change Detection, a Case Study of Kisumu East. Nairobi: University of Nairobi.
65
Ochieng’-SpringerS. (2022). Governance and public administration during the COVID-19 pandemic: issues and experiences in Kenya’s health system. Politikon49 (1), 1–20. 10.1080/02589346.2021.2008091
66
OdhiamboF. O. (2019). Assessing the predictors of lived poverty in Kenya: a secondary analysis of the Afrobarometer survey 2016. J. Asian Afr. Stud.54 (3), 452–464. 10.1177/0021909618822668
67
OlekuS. R. (2024). Assessing the Impacts of Climate Change on Water Resources and Pastoralists Livelihoods in Kajiado West Sub-county, Kenya. Nairobi: University of Nairobi, 97.
68
OluchiriS. O. (2025). Urban flooding in the cities of Kisumu, Mombasa, and Nairobi, Kenya: causes, vulnerability factors, and management. Afr. J. Empir. Res.6 (1), 342–351. 10.51867/ajernet.6.1.29
69
OmbogoL. (2016). Rainfall Trends and flooding in the Sondu miriu river basin. erepository.uonbi23, 1–30.
70
OpolotE. (2013). Application of Remote Sensing and Geographical Information Systems in Flood Management: A Review.
71
OthooC. O. (2021). The impact of climate extremes (floods) on sanitation infrastracture in unplanned settlements in Kisumu city, Kenya. Nairobi: University of Nairobi.
72
ParkK.LeeE. H. (2024). Urban flood vulnerability analysis and prediction based on the land use using Deep Neural Network. Int. J. Disaster Risk Reduct.101, 104231. 10.1016/j.ijdrr.2023.104231
73
PaudelS.YuanF. (2012). Assessing landscape changes and dynamics using patch analysis and GIS modeling. Int. J. Appl. Earth Observ. Geoinformat.16, 66–76. 10.1016/j.jag.2011.12.003
74
PaulP. K. A. P. S.BhuimaliA.TiwaryK. S.SaavedraR.AremuB. (2020). “Geo information systems and remote sensing: applications in environmental systems and management,” in GIS in Env Reviewed, New York, NY: Routledge77.
75
PeramunaP.NeluwalaN.WijesundaraK.DeSilvaS.VenkatesanS.DissanayakeP. (2025). Enhancing 2D hydrodynamic flood model predictions in data-scarce regions through integration of multiple terrain datasets. J. Hydrol.648, 132343. 10.1016/j.jhydrol.2024.132343
76
PlateE. J. (2002). Flood risk and flood management. J. Hydrol.267 (1-2), 2–11. 10.1016/s0022-1694(02)00135-x
77
PradhanB. (2010). Flood susceptible mapping and risk area delineation using logistic regression, GIS and remote sensing. J. Spatial Hydrol.9 (2), 1–20.
78
PrasadP.LovesonV. J.DasB.KothaM. (2022). Novel ensemble machine learning models in flood susceptibility mapping. Geocarto Int.37 (16), 4571–4593. 10.1080/10106049.2021.1892209
79
PuttinaovaratS.HorkaewP. (2020). Flood forecasting system based on integrated big and crowdsource data by using machine learning techniques. IEEE Access8, 5885–5905. 10.1109/access.2019.2963819
80
RahmatiO.ZeinivandH.BesharatM. (2016). Flood hazard zoning in Yasooj region, Iran, using GIS and multi-criteria decision analysis. Geomatics, Nat. Hazards Risk7 (3), 1000–1017. 10.1080/19475705.2015.1045043
81
SadiqR.AkhtarZ.ImranM.OfliF. (2022). Integrating remote sensing and social sensing for flood mapping. Remote Sens. Appl. Soc. Environ.25, 100697. 10.1016/j.rsase.2022.100697
82
SchoppaL.DisseM.BachmairS. (2020). Evaluating the performance of random forest for large-scale flood discharge simulation. J. Hydrol.590, 125531. 10.1016/j.jhydrol.2020.125531
83
SenguptaS. (2024). IoT-Based flood detection and management systems in urban areas. Risk Assess. Manag. Decis.1 (2), 301–313. 10.1007/978-3-031-55567-3_15
84
ShahM. A. R.RahmanA.ChowdhuryS. H. (2018). Challenges for achieving sustainable flood risk management. J. Flood Risk Manag.11, S352–S358. 10.1111/jfr3.12211
85
SharmaT. P. P.ZhangJ.KojuU. A.ZhangS.BaiY.SuwalM. K. (2019). Review of flood disaster studies in Nepal: a remote sensing perspective. Int. J. Dis. Risk Reduct.34, 18–27. 10.1016/j.ijdrr.2018.11.022
86
ShastryA.DurandM. (2019). Improved DEM development for flood inundation simulation. IEEE Trans. Geoscience Remote Sens.57 (8), 1–12. 10.1109/TGRS.2019.2914995
87
Sheikh KumranA. N. S.NurP. N. M.UmberN.NurA. A. (2020). A review on the application of remote sensing and geographic information system in flood crisis management. J. Crit. Rev.7 (16), 1–10.
88
ShenY.ZhuZ.ZhouQ.JiangC. (2024). An improved dynamic bidirectional coupled hydrologic–hydrodynamic model for efficient flood inundation prediction. Nat. Hazards Earth Syst. Sci.24 (7), 2315–2330. 10.5194/nhess-24-2315-2024
89
ShiH.LiT.LiuR.ChenJ.LiJ.ZhangA.et al (2015). A service-oriented architecture for ensemble flood forecast from numerical weather prediction. J. Hydrol.527, 933–942. 10.1016/j.jhydrol.2015.05.056
90
SinghaC.RanaV. K.PhamQ. B.NguyenD. C.ŁupikaszaE. (2024). Integrating machine learning and geospatial data analysis for comprehensive flood hazard assessment. Environ. Sci. Pollut. Res.31 (35), 48497–48522. 10.1007/s11356-024-34286-7
91
StefanidisS.StathisD. (2013). Assessment of flood hazard based on natural and anthropogenic factors using analytic hierarchy process (AHP). Nat. Hazards68 (2), 569–585. 10.1007/s11069-013-0639-5
92
TangY.SunY.HanZ.SoomroS. e. h.WuQ.TanB.et al (2023). Flood forecasting based on machine learning pattern recognition and dynamic migration of parameters. J. Hydrol. Regional Stud.47, 101406. 10.1016/j.ejrh.2023.101406
93
TanimA. H.McRaeC. B.Tavakol-DavaniH.GoharianE. (2022). Flood detection in urban areas using satellite imagery and machine learning. Water14 (7), 1140. 10.3390/w14071140
94
Thomas Steven SavageJ.PianosiF.BatesP.FreerJ.WagenerT. (2016). Quantifying the importance of spatial resolution and other factors through global sensitivity analysis of a flood inundation model. Water Resour. Res.52 (11), 9146–9163. 10.1002/2015wr018198
95
Tien Bui, D.Khosravi, K.ShahabiH.DaggupatiP.AdamowskiJ. F.MelesseA. M.et al (2019). Flood Spatial Modeling in Northern Iran Using Remote Sensing and Gis: A Comparison Between Evidential Belief Functions and its Ensemble with a Multivariate. Remote Sens.11 (13), 1589. 10.3390/rs11131589
96
Tien BuiDHoangN‐D.Martínez‐ÁlvarezF.Thi NgoP‐T.Viet HoaP.Dat PhamT.et al (2020). A novel deep learning neural network approach for predicting flash flood susceptibility: a case study at a high frequency tropical storm area. ScienceDirect701, 1–15. 10.1016/j.scitotenv.2020.139985
97
TsakirisG. (2014). Flood risk assessment: concepts, modelling, applications. Nat. Hazards Earth Syst. Sci.14 (5), 1361–1369. 10.5194/nhess-14-1361-2014
98
TsanakasK.KarymbalisE.GrivaD.ValkanouK.BatzakisD. V.VassilakisE.et al (2025). The geomorphology of Greece. J. Maps21 (1), 2540555. 10.1080/17445647.2025.2540555
99
VeerakachenW.RaksapatcharawongM. (2015). Rainfall estimation for real time flood monitoring using geostationary meteorological satellite data. Adv. Space Res.56 (6), 1139–1145. 10.1016/j.asr.2015.06.016
100
VilanculosA. C. F. (2015). The Use of Hydrological Information to Improve Flood management-integrated Hydrological Modelling of the Zambezi River Basin. Grahamstown: Rhodes University, 273.
101
WuR. E.DuanQ.WoodA. W. (2020a). Ensemble flood forecasting: Current status and future opportunities. London, Wiley.
102
WahbaM.EssamR.El-RawyM.Al-ArifiN.AbdallaF.ElsadekW. M. (2024). Forecasting of flash flood susceptibility mapping using random forest regression model and geographic information systems. Heliyon10 (13), e33982. 10.1016/j.heliyon.2024.e33982
103
WangX. H. X.-W. (2018). A Review on Applications of Remote Sensing and Geographic Information Systems (GIS) in Water Resources and Flood Risk Management. New York, NY: ProQuest.
104
WangZ.LaiC.ChenX.YangB.ZhaoS.BaiX. (2015). Flood hazard risk assessment model based on random forest. J. Hydrol.527, 1130–1141. 10.1016/j.jhydrol.2015.06.008
105
WangH.MengY.XuH.WangH.GuanX.LiuY.et al (2024). Prediction of flood risk levels of urban flooded points though using machine learning with unbalanced data. J. Hydrol.630, 130742. 10.1016/j.jhydrol.2024.130742
106
WannewitzM.GarschagenM. (2021). Mapping the adaptation solution space–lessons from Jakarta. Nat. Hazards Earth Syst. Sci.21 (11), 3285–3322. 10.5194/nhess-21-3285-2021
107
WanzalaM. A. (2022). Improving Flood Modelling and Forecasting in Kenya. Reading: University of Reading.
108
WanzalaM. A.StephensE. M.ClokeH. L.FicchiA. (2025). Hydrological model preselection with a filter sequence for the national flood forecasting system in Kenya. J. Flood Risk Manag.18 (1), e12846. 10.1111/jfr3.12846
109
WaskoC.WestraS.NathanR.OrrH. G.VillariniG.Villalobos HerreraR.et al (2021). Incorporating climate change in flood estimation guidance. Philosophical Trans. R. Soc. A379 (2195), 20190548. 10.1098/rsta.2019.0548
110
WingO. E.PinterN.BatesP. D.KouskyC. (2020). New insights into US flood vulnerability revealed from flood insurance big data. Nat. Commun.11 (1), 1444. 10.1038/s41467-020-15264-2
111
WinsemiusH. C.AertsJ.van BeekL.BierkensM.BouwmanA.JongmanB.et al (2016). Global drivers of future river flood risk. Nat. Clim. Change6 (4), 381–385. 10.1038/nclimate2893
112
WuH.AdlerR. F.HongY.TianY.PolicelliF. (2012). Evaluation of global flood detection using satellite-based rainfall and a hydrologic model. J. Hydrometeorol.13 (4), 1268–1284. 10.1175/jhm-d-11-087.1
113
WuW.EmertonR.DuanQ.WoodA. W.WetterhallF.RobertsonD. E. (2020b). Ensemble flood forecasting: current status and future opportunities. Wiley Interdiscip. Rev. Water7 (3), e1432. 10.1002/wat2.1432
114
YalcinE.AkyurekZ. (2020). The effect of DEM resolution and roughness on flood inundation modeling. J. Flood Risk Manag.13 (2), 1–15. 10.1111/jfr3.12613
115
YariyanP.JanizadehS.Van PhongT.NguyenH. D.CostacheR.Van LeH.et al (2020). Improvement of best first decision trees using bagging and dagging ensembles for flood probability mapping. Water Resour. Manag.34 (9), 3037–3053. 10.1007/s11269-020-02603-7
116
ZhouS.WangY.LiZ.ChangJ.GuoA. (2021). Quantifying the uncertainty interaction between the model input and structure on hydrological processes. Water Resour. Manag.35 (12), 3915–3935. 10.1007/s11269-021-02883-7
117
ZotouI.BellosV.GkoumaA.KarathanassiV.TsihrintzisV. A. (2020). Using Sentinel-1 imagery to assess predictive performance of a hydraulic model. Water Resour. Manag.34 (14), 4415–4430. 10.1007/s11269-020-02592-7
Summary
Keywords
ensemble methods for classification, flood risk modeling, Lake Victoria (East Africa), machine learning-based classification, remote sensing, GIS, uncertainty quantification
Citation
Auma H, Gebreslasie M and Osio A (2026) Flood risk modeling for Kisumu County, Kenya: integrating multi-source geospatial data and machine learning. Front. Environ. Sci. 14:1829199. doi: 10.3389/fenvs.2026.1829199
Received
12 March 2026
Revised
05 May 2026
Accepted
06 May 2026
Published
11 August 2026
Volume
14 - 2026
Edited by
Pasquale Imperatore, National Research Council Naples (CNR), Italy
Reviewed by
Alessio Di Simone, University of Naples Federico II, Italy
Muhammad Habib Ullah, National Research Council (CNR), Italy
Olusegun Adeaga, University of Lagos, Nigeria
Updates
Copyright
© 2026 Auma, Gebreslasie and Osio.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Herine Auma, 225191174@stu.ukzn.ac.za
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.