ORIGINAL RESEARCH article

Front. Earth Sci., 31 October 2025

Sec. Geohazards and Georisks

Volume 13 - 2025 | https://doi.org/10.3389/feart.2025.1702688

Spatial consistency assessment and landslide susceptibility prediction optimization

  • 1. Nanjing Changtian Surveying and Mapping Technology Co., Ltd., Nanjing, China

  • 2. School of Geography Science and Geomatics Engineering, Suzhou University of Science and Technology, Suzhou, China

Abstract

Currently, although various landslide susceptibility models can achieve high prediction accuracy, their results have significant differences in spatial distribution, resulting in high prediction uncertainty, which poses a challenge to optimizing assessment methods applicable to such complex geological hazards. In order to reduce uncertainty, this study proposes a machine learning ensemble modeling method that combines spatial consistency analysis. Taking Ruijin City in Jiangxi Province as the research area, based on the selection of 12 influencing factors and hyperparameter optimization, three algorithms including XGBOOST, Random Forest (RF), and Support Vector Machine (SVM) were used to generate landslide susceptibility maps. All models performed well, with AUC values ranging from 0.84 to 0.93. However, spatial consistency analysis shows that the spatial correlation between maps between models is only 0.78 to 0.84, indicating that although the prediction accuracy is high, there is still significant spatial heterogeneity and uncertainty. Therefore, a logistic regression (LR) fusion model based on historical landslides was constructed. Use the compilation results as the dependent variable and the results of the three models as the independent variables. The results indicate that XGBOOST contributes the most, followed by RF and SVM. By integrating the three prediction results, a comprehensive vulnerability map was finally obtained, which was superior to the single model in terms of spatial consistency (correlation coefficient 0.87–0.91) and prediction accuracy (AUC = 0.95). This research framework effectively reduces the uncertainty of landslide prediction and improves the reliability and accuracy of evaluation results.

1 Introduction

Landslides are highly destructive natural disasters that threaten human safety, socioeconomic stability, and ecological sustainability (; ; ). Their abruptness and uncertainty make timely landslide information crucial for risk management (). Consequently, landslide susceptibility assessment has become a key tool for identifying potential hazards and supporting disaster prevention planning (; ; ). Advances in remote sensing have significantly improved the availability of high spatio-temporal resolution Earth observation data, enhancing landslide susceptibility mapping. High-resolution satellite imagery enables the extraction of key environmental parameters—including topography, vegetation cover, geological structure, and hydrology—which are vital for landslide indication. Additionally, remote sensing plays a pivotal role in identifying and compiling historical landslide inventories, providing reliable data for disaster records (). As data quality continues to improve, methodological choices increasingly determine assessment reliability (). Researchers have developed various models using GIS and remote sensing, such as weight of evidence, logistic regression, analytic hierarchy process, and evidential belief function. Recently, machine learning applications have grown substantially (). Early introduced methods like K-nearest neighbor (KNN) have been followed by widely adopted algorithms including Logistic Regression (LR) and Support Vector Machine (SVM), valued for their adaptability (). Advanced techniques like extreme gradient boosting (e.g., XGBoost) and ensemble learners such as Random Forests (RF) have demonstrated superior predictive performance (). However, differences in model selection, data sources, and human judgment often introduce substantial uncertainties in landslide susceptibility assessments (). Since high-prediction landslide maps are critical for disaster decision-making, they must undergo rigorous validation before use (). Currently, two main challenges remain: how to accurately evaluate susceptibility maps, and how to identify the optimal method combination to enhance efficiency. Conventional verification involves susceptibility simulation and result-field comparison, demanding reliability, robustness, and predictive capability from the method (; ). Notably, even when models perform similarly on test datasets, their spatial predictions may still vary significantly ().

While different machine learning algorithms have been employed for landslide susceptibility mapping, the pixel-level consistency among these methods is not well studied. The spatial heterogeneity they produce further elevates assessment uncertainty (; ; ). Hence, this study introduces an integrated modeling approach to minimize prediction uncertainty. This is achieved by evaluating the consistency among three machine learning results and fusing them into a comprehensive susceptibility map. Our case study is Ruijin City, Jiangxi Province, where complex geological conditions, high rainfall, and documented landslide events make it a representative area. The method’s effectiveness and applicability will be further verified through field data and historical landslide records.

2 Materials and methods

This study consisted of three main phases. In the first stage, three machine learning algorithms, including XGBOOST, Random Forest (RF) and Support Vector Machine (SVM), were used to generate landslide susceptibility maps in the study area. In the second stage, the consistency of the spatial prediction patterns of landslide probability maps obtained by different methods was evaluated based on pixel-by-pixel correlation analysis. In the third stage, the output results of the three models are fused to synthesize a comprehensive landslide susceptibility regionalization map.

2.1 Study area

This study focuses on Ruijin City, Jiangxi Province (25°30′–26°20′N, 115°42′–116°22′E), a hilly and mountainous region prone to landslides. The area experiences high annual rainfall (>1,600 mm) and frequent human activities, which collectively contribute to slope instability (Figure 1). Historical landslides, triggered by heavy rainfall and construction, have repeatedly damaged infrastructure and threatened public safety. Therefore, landslide susceptibility assessment is crucial for disaster prevention and spatial planning in Ruijin.

FIGURE 1

2.2 Landslide cataloging and mapping

Landslide inventory mapping is a fundamental step in landslide susceptibility assessment (). This study utilized the official landslide inventory map of Ruijin City, produced by the Jiangxi Provincial Geological Bureau. This map integrates historical landslide records from various sources, verified through GPS and field surveys. Additionally, our field investigations provided detailed reports on landslide movement types, distribution patterns, damages, displacement, materials, and triggers. The study is based on 370 identified landslide sites. Since landslide susceptibility mapping is a binary classification task, non-landslide samples are equally critical. Following (; ), who outlined three methods for selecting non-landslide samples, this study adopted the second approach: randomly selecting points from areas with no landslide history. Using ArcGIS, we generated 10 sets of non-landslide data, each containing 370 points. The entire dataset (landslide and non-landslide) was then divided, with 70% allocated for model training and 30% for testing.

2.3 Landslide influencing factors

The effectiveness of landslide susceptibility mapping largely depends on the selection of influencing factors (; ; ). In this study, the following principles were followed in selecting these factors (; ; ): (1) the factor must have a known mechanical or statistical association with landslide occurrence; (2) the factor must be quantifiable and spatially mappable; (3) redundancy among factors should be minimized to reduce multicollinearity issues in the model; and (4) the factor must align with the geomorphological and geological characteristics of the study area (Table 1).

TABLE 1

Environmental factorsValuesNumber of grids in the whole areaGrid scale/%Landslide gridLandslide grid scale/%FR
Elevation(m)139.7–250.9836,74530.41917346.7571.152
250.9–335.3796,48228.95612132.7031.175
335.3–423.5578,69121.0384913.2430.796
423.5–538.6332,05612.072205.4050.650
538.6–695.9147,9305.37841.0810.159
695.9–1,117.858,7872.13730.8110.376
Slope(°)0–4.4685,21824.9114111.0810.260
4.4–8.8643,53523.39512533.7840.986
8.8–13.2608,75522.13111330.5411.276
13.2–17.9446,52016.2335615.1351.945
17.9–28.7344,70312.532349.1890.632
28.7–51.221,9600.79810.2700.398
Aspect−1155,9405.669277.2970
0–22.5297,92410.831318.3780.994
22.5–67.5354,47912.8876215.7570.954
67.5–112.5359,79113.0804812.9731.301
112.5–157.5332,83012.0995414.5951.198
157.5–202.5332,14312.0754211.3511.160
202.5–247.5378,01113.7424812.9731.086
247.5–292.5370,19513.4583810.2700.792
292.5–337.5169,3786.158205.4050.716
Profile curvature0–2.029884,49932.1569826.4860.596
2.029–4.057773,41628.11712533.7841.072
4.057–6.324561,55120.4156918.6491.126
6.324–8.949324,37611.7935414.5951.213
8.949–14.529187,6476.822236.2160.823
14.529–30.42819,2020.69810.2701.829
Plan curvature0–13.422651,67723.69111130.0001.246
13.422–24.927625,67522.7469926.7571.417
24.927–37.710471,54417.1435815.6761.434
37.710–52.091354,66612.8944211.3510.932
52.091–67.749301,69610.968246.4860.657
67.749–81.491345,43312.558369.7290.852
Topographic relief0–6.022651,45023.683359.4590.476
6.022–12.420721,23626.22014840.0000.566
12.420–18.819641,93823.33710428.1081.163
18.819–22.969293,59710.6744311.6222.575
22.969–35.379385,79914.0263810.2700.431
35.379–95.97572,8552.64930.8110.367
LithologyMetamorphic rock121858444.30110829.1891.301
Magmatic rock503,74818.314277.2971.611
Clastic rock899,36332.696195.1350.209
Carbonatite128,9964.68921658.3780.659
NDVI−0.054–0.00668,0982.47651.3510.192
0.006–0.018299,11510.874349.1890.803
0.018–0.025580,37321.0995615.1351.009
0.025–0.033848,42030.84314639.4590.955
0.033–0.042635,13223.0898723.5141.075
0.042–0.098315,48811.4694211.3511.141
NDBI−0.650∼-0.38974,9632.725133.5131.101
−0.389∼-0.318234,6328.529287.5680.928
−0.318∼-0.267428,67415.5845815.6761.418
−0.267∼-0.219699,58125.4339224.8651.233
−0.219∼-0.173803,44529.20911029.7290.901
−0.173∼-0.050505,33118.3716918.6490.729
MNDWI−0.035–0.110365,88213.3014812.9731.374
0.110–0.164773,62128.12511831.8921.221
0.164–0.217772,21228.0739425.4050.952
0.217–0.276492,15817.8926718.1081.174
0.276–0.352256,7189.333297.8381.082
0.352–0.64386,0353.128133.5140.708
Distance to river(m)<150155,2125.6424712.7032.586
150–30055,8082.02971.8911.689
300–450279,11410.14711831.8920.672
>450227411682.67419853.5140.497
Distance to roads(m)<150265,20631.43111230.2700.963
150–300366,47928.1349024.3243.139
300–450337,20121.7898222.1621.663
450–600599,35112.2594512.1620.558
600–800773,8726.052349.1890.327
>800408,5820.33571.8910.127

Frequency ratio and related description of each influencing factor.

Based on the aforementioned principles, existing literature, and professional understanding of the study area, an initial set of factors encompassing topography, geology, hydrology, vegetation cover, and human activities was selected. It should be specifically noted that although rainfall is a key dynamic trigger for landslides, it was not directly incorporated into the model due to its limited spatial variability at the regional scale and the study’s focus on assessing long-term static susceptibility. Similarly, high-resolution soil moisture data and detailed construction activity data were excluded because they were difficult to systematically obtain and standardize across the study area. As an alternative, remotely sensed indices (e.g., NDBI) and distance-based factors (e.g., proximity to roads) were used to indirectly yet effectively represent the intensity and distribution of human activities. Ultimately, 12 influencing factors were identified for modeling (Table 1). All factors were derived from 30-m spatial resolution ALOS DEM and Landsat satellite imagery, processed via the Google Earth Engine platform. The following (Figure 2) provides a detailed description of each factor category:

FIGURE 2

Topographic Factors: These include elevation, slope, aspect, plan curvature, profile curvature, and topographic relief. Collectively, they govern slope morphology, stress distribution, and surface drainage conditions, forming the intrinsic basis for landslide occurrence (; ; ). For instance, slope directly influences gravity-driven shear stress, while curvature relates to the convergence or divergence of surface materials.

Geological Factor: Lithology. Variations in the strength and permeability of different rock and soil types directly control slope stability and failure mechanisms.

Hydrological Factor: Distance to rivers. Riverbank erosion is a significant external force triggering landslides. A distance-to-river map was generated using the Euclidean distance algorithm.

Vegetation and Surface Cover Factors: This study incorporated three complementary remote sensing indices to comprehensively characterize the surface environment:

Normalized Difference Vegetation Index (NDVI): Quantifies vegetation density. Dense vegetation enhances soil shear strength through root reinforcement, while sparse vegetation areas are more prone to shallow landslides.

Modified Normalized Difference Water Index (MNDWI): Accurately extracts water bodies. Areas near water are not only threatened by lateral erosion but are also affected by dynamic groundwater levels that influence slope stability.

Normalized Difference Built-up Index (NDBI): Identifies built-up areas. This index effectively reflects the alteration and disturbance of natural slopes by human activities (e.g., land excavation, engineering loads). The combined use of NDBI, NDVI, and MNDWI holistically captures the spatial pattern of “vegetation-water-built-up” areas, providing a more integrated perspective on how human-environment interactions influence landslide risk.

Human Activity Factor: Distance to roads. Road construction often involves large-scale cutting and filling, significantly disrupting the natural equilibrium of slopes. A distance-to-roads map was generated using the Euclidean distance algorithm.

Prior to modeling, the variance inflation factor (VIF) for all 12 factors was calculated using the R platform to assess multicollinearity. As shown in Table 2, all factors had VIF values below 2.8, indicating no severe multicollinearity issues, thus confirming their suitability for subsequent modeling analysis.

TABLE 2

Impact factorsVIF
Elevation1.418736
Slope1.422683
Aspect1.135443
Curvature of the plane1.574954
Profile curvature1.574954
Formation lithology2.37445
Relief of topography1.695373
NDVI2.351282
NDBI2.625348
MNDWI1.979472
Distance from road1.658247
Distance from river1.711446

Estimated variance inflation factors of landslide impact condition factors.

2.4 Multicollinearity analysis

Before modeling landslide susceptibility, it is necessary to test the correlation between various potential hazard factors to identify possible multicollinearity problems (). To this end, with the help of the R language platform, this study calculated the variance inflation factor (VIF) for each of the selected 12 landslide impact factors, which is often used to evaluate the degree of collinearity between independent variables. A VIF value of more than 5 for a variable is generally considered to indicate significant multicollinearity. As shown in Table 2, the VIF values of all the factors in this study were below 2.8, which indicated that there was no significant collinearity problem between these variables and could be used for subsequent modeling analyses.

3 Modeling landslide susceptibility

3.1 Data preprocessing

In the GIS platform, the corresponding values of 12 influencing factors were extracted according to the spatial distribution of landslide sites and non-landslide sites. These factors included 8 continuous variables and 4 discrete variables. Discrete categorical variables were converted to composite binary feature forms, generating dummy variables that were consistent with the number of categories (). Specifically, one-hot encoding method is used for processing. For example, geological types contain 11 categories, and if a location belongs to one of these categories, this category is coded as 1, and the other categories are marked as 0. Other discrete variables are also coded in the same way. To further improve modeling efficiency, all continuous variables were standardized: the mean and standard deviation of each variable were calculated, and each observation was divided by the standard deviation after subtracting the mean. This process not only unifies the dimensions, but also helps to narrow the parameter search range of the optimization algorithm, thus speeding up the model training process.

3.2 Hyperparameter optimization

Hyperparameter optimization systematically evaluates different parameter combinations to identify the optimal configuration, thereby improving the prediction accuracy of machine learning models (). In this study, we implemented hyperparameter tuning using grid search with five-fold cross-validation on the training set. The resulting optimal hyperparameters were used for final model training and testing. For instance, the Random Forest model achieved best performance with 500 decision trees (Table 3). This sufficient number of trees helps integrate diverse predictions, mitigate the impact of individual tree randomness on susceptibility mapping, and enhance overall robustness. All hyperparameter settings, search ranges, and final values are documented in Table 2.

TABLE 3

ClassierHyperparameterSearch rangeOptimal value
RFNumber of estimators200, 300, 400, 500500
Maximum featuresAuto, square root, logarithm (base = 2)Auto
Maximum depth10, 12, 14, 16, 18, 20, 22, 24, 26, 2810
CriterionGini, entropyEntropy
SVMC value10–3, 10–2, 10–1, 1, 10, 102, 103103
KernelPolynomial, radial basis function, sigmoidRadial basis function
Gamma10–4,10–3, 10–2,10–110–4
XGBoostn_estimators100, 200, 300, 400, 500400
max_depth3, 4, 5, 6, 7, 8, 9, 106
learning_rate0.01, 0.05, 0.1, 0.2, 0.30.1
subsample0.6, 0.7, 0.8, 0.9, 1.00.8
colsample_bytree0.6, 0.7, 0.8, 0.9, 1.00.8
gamma0, 0.1, 0.2, 0.3, 0.4, 0.50.2
reg_alpha0, 0.1, 0.5, 1.0, 2.00.5
reg_lambda0.5, 1.0, 1.5, 2.0, 2.51.0

Hyperparameters, search ranges, and optimal values of landslide susceptibility models based on machine learning.

3.3 Machine learning models

3.3.1 RF

As an ensemble learning algorithm, Random Forest (RF) has shown good performance in landslide susceptibility prediction in recent years (; ; ). By constructing multiple decision trees and performing ensemble voting, the model can effectively deal with high-dimensional nonlinear data, and has excellent generalization ability and anti-overfitting characteristics. In the application of landslide prediction, RF model can comprehensively deal with a variety of environmental factors (such as elevation, slope, lithology, rainfall, etc.), generate diversified decision trees by Bootstrap sampling and random feature selection, and finally output the landslide potential of each region in the form of probability. The results show that RF can not only evaluate the importance of each influencing factor, but also stably generate high-precision prediction results in complex geographic environments, which provides a reliable basis for regional landslide risk management and land planning.

3.3.2 XGBOOST

XGBoost (Extreme Gradient Boosting) is a kind of efficient Gradient to promote integration algorithm, susceptibility in landslide prediction shows good performance (; ). In this model, multiple decision trees are built iteratively, and each tree is dedicated to correcting the prediction error of the previous round, so as to gradually improve the overall prediction accuracy. XGBOOST can automatically deal with the complex interaction between features, and has good compatibility for continuous and categorical variables. It is suitable for integrating multi-source environmental factors (such as elevation, slope, lithology, rainfall, land cover, etc.) to assess landslide sensitivity. Its key advantages include regularization to prevent overfitting, built-in cross-validation, and parallel computing to accelerate the training process. In landslide prediction applications, XGBOOST can not only output the probability of landslide occurrence for each spatial unit, but also provide feature importance ranking to help identify key impact factors and enhance the interpretability of the model. Studies show that XGBoost is usually superior to traditional machine learning models (such as logistic regression or single decision tree) in dealing with high-dimensional geospatial data, which can more accurately depict the nonlinear relationship between landslides and driving factors, and provide high-precision prediction basis for regional landslide risk management and land planning.

3.3.3 SVM

SVM (Support Vector Machine, SVM) in landslide prone forecasts are widely used in dealing with high-dimensional nonlinear classification (; ; ). The model by looking for the optimal hyperplane and maximizing in the feature space between the positive and negative samples (with the landslide) the classification of the interval, which have good generalization ability. When dealing with landslide prediction tasks, SVM can comprehensively utilize multiple environmental factors such as topography, geology, hydrologic and human activities, and map nonlinear relationships through kernel functions (such as RBF kernel) to effectively depict the complex interactions between landslide occurrence and influencing factors. The results show that SVM can maintain high classification accuracy in the case of limited sample size, and its robustness to noise data and clear mathematical derivation mechanism make it a reliable and interpretable modeling tool in landslide hazard assessment.

3.3.4 Performance evaluation methods

Model performance was evaluated using the receiver operating characteristic (ROC) curve and its area under the curve (AUC) as the primary criteria. The ROC curve was plotted using 30% of the test data, with the false positive rate (1 - specificity) on the x-axis and sensitivity (recall) on the y-axis. Sensitivity measures the model’s ability to correctly identify landslides, calculated as the proportion of true positives among all actual landslides. Specificity, the proportion of true negatives among all non-landslide samples, indicates how well the model excludes non-landslide areas. AUC values were interpreted as: 0.5–0.6 (poor), 0.6–0.7 (fair), 0.7–0.8 (good), 0.8–0.9 (excellent), and 0.9–1.0 (outstanding). Additional metrics derived from the confusion matrix—including accuracy, precision, recall, and F1-score—provided further validation of model performance.

3.3.5 Spatial consistency analysis and multi-model ensemble optimization for landslide susceptibility prediction

In order to analyze the consistency of the prediction results of different models, this study evaluated the consistency of landslide susceptibility maps generated by multiple machine learning algorithms in spatial distribution by pixel-by-pixel comparison. For four kinds of models of six possible combination of two, the Pearson correlation coefficient are calculated respectively. The coefficient is defined as the ratio of the covariance between the predictions of the two models and the product of their respective standard deviations, with values ranging from −1 to +1:0 for no correlation, less than ±0.29 for low agreement, ±0.30 to ±0.49 for moderate agreement, ±0.50 to ±1 (excluding ±1) for high agreement, and ±1 for perfect agreement. After complete the spatial consistency analysis, the further integration of the four machine learning model output, to generate an optimized integrated landslide prone figure. Integrated methods using logistic regression (LR) model, with dual landslide logging data (that is, the point with the landslide points) as the dependent variable, in four different forecast results as the independent variable of the model. The regression coefficients of each model output were obtained by fitting, and the landslide occurrence probability (P) of each pixel was calculated based on the Equation 1 in the Geographic Information System (GIS) platform, so as to obtain the comprehensive landslide probability distribution map of the study area.

Where, is a linear combination of independent variables, and its calculation formula is as follows:

Where is the model intercept; is the regression coefficient of the independent variable; Represents n independent variables. The ROC curve was drawn using 30% of the test data to verify the final obtained comprehensive model.

4 Results

4.1 Landslide prediction

Figure 3 shows the landslide susceptibility distribution maps generated based on three different machine learning methods in Ruijin City. In the GIS platform, all the output probability maps were divided into five susceptibility levels by using the natural breakpoint method: (1) very low susceptibility (0-0.1), (2) low susceptibility (0.11-0.3), (3) medium susceptibility (0.31-0.5), (4) high susceptibility (0.51-0.85), and (5) very high susceptibility (0.86-1). It can be seen from Figure 3 that there is a significant difference in the proportion of areas predicted by each model for high susceptibility regions. The total area of “high” and “extremely high” areas identified by XGBoost model accounted for the largest proportion, reaching 38.2%. However, the proportion of these three types of areas in the results obtained by the SVM model was the lowest, only 20.2%.

FIGURE 3

4.2 Model performance evaluation

To assess the performance of different landslide susceptibility model, this study randomly selected 30% of the data as a test set, and on the basis of constructing the performance comparison matrix (see Table 3). All the evaluation indicators showed that all the models showed high prediction accuracy. In terms of the overall classification accuracy, XGBOOST method performed the best, reaching 90.16%. This was followed by RF (88.39%) and SVM (84.28%). However, as a comprehensive performance index, the overall accuracy is difficult to identify the classification bias in specific categories. Therefore, in order to further test the consistency of the model in the discrimination of landslide points and non-landslide points, this paper additionally calculated the precision, recall and F1 score (Table 4). From the results, XGBOOST keeps leading in all indicators, and RF also performs stably and closely behind. Meanwhile, in terms of the area under the receiver operating characteristic curve (AUC), XGBOOST also ranked first with 0.930, while RF and SVM were 0.885 and 0.845, respectively (Figure 4). The excellent performance of RF and SVM models can be attributed to their ability to effectively capture the complex nonlinear relationship between regional geographical characteristics and landslide occurrence.

TABLE 4

ModelAccuracyPrecisionF1Recall
Non-landslideLandslideNon-landslideLandslideNon-landslideLandslide
SVM0.84280.83590.83590.82380.82380.81170.8117
RF0.88390.87630.87630.86580.86580.85420.8542
XGBoost0.90160.89120.89120.88830.88830.87560.8756

Performance evaluation indicators of landslide susceptibility models.

FIGURE 4

4.3 Spatial consistency of different methods

In this study, a correlation matrix (Figure 5) was constructed by comparing the landslide occurrence probabilities obtained by different methods pixel by pixel to analyze the level of agreement between various landslide susceptibility models. The results showed that although the AUC values of different models were similar, the spatial consistency of landslide susceptibility distribution maps (LSM) generated by them was still significantly different. In general, the correlation coefficients between the models ranged from 0.78 to 0.84. Among them, the combination of XGBOOST and RF shows the highest consistency, while the combination of SVM and RF shows the lowest consistency. An integrated landslide susceptibility map (LSM) was generated by substituting the logistic regression (LR) coefficients and intercepts of the three models into Equation 2. The resulting LSM was divided into five vulnerability levels following the natural breakpoint classification method. Compared with the prediction results of a single model, the area under the ROC curve (AUC) reached the highest 0.9537 (Figure 4). In the study area, the total area classified as “high” and “extremely high” risk level accounted for about 27.3%. From the perspective of spatial consistency, the comprehensive landslide susceptibility map showed high correlation with the results of each single model, and the correlation coefficients ranged from 0.87 to 0.91 (Figure 5).

FIGURE 5

4.4 Comprehensive landslide susceptibility mapping

Different machine learning methods often show significant spatial heterogeneity in landslide prediction. In order to reduce the uncertainty caused by a single model, this study generates a comprehensive landslide susceptibility map by integrating the output results of multiple algorithms. A regression-based fusion strategy was used to construct a multiple Logistic regression (LR) model with the binary landslide cataloging data as the dependent variable and the prediction results of the three models as the independent variables. The results of the regression model showed that the regression coefficients of SVM, Random Forest (RF) and XGBOOST methods were statistically significant (P < 0.05). The overall goodness-of-fit of the model was high, and the coefficient of determination (R2) reached 0.80. The regression coefficients showed that the XGBOOST model had the strongest consistency with the real landslide distribution, followed by RF and SVM. This ranking is consistent with the performance of each model as reflected by the AUC value when predicting the landslide separately (Figure 4), indicating that the regression weight effectively reflects the contribution of different models in the ensemble.

There were obvious spatial differences in the landslide susceptibility distribution of different township units in Ruijin city. The landslide susceptibility level was relatively high in Xifang Town and Yeping town, and the area of “high” and “extremely high” susceptibility grade in these two towns accounted for more than 7% of the total area of the corresponding grade in the study area. Town and cortex phellodendri conventions in addition, as 6.5% of the area were classified as high rock landslide. It is worth noting that the built-up areas and infrastructure coverage of some villages and towns located in hilly areas significantly overlap with the areas with high susceptibility to landslides. Surveys in recent years have shown that with the intensification of urban and rural construction activities, local slope excavation, vegetation damage, and hydrological changes have further increased the risk of landslide hazards in these areas, posing potential threats to residents’ safety and engineering facilities.

This study further integration the prediction results of three kinds of machine learning algorithm, and generate a map on integrated landslide susceptibility (LSM). The results showed that the combined results were superior to any single model in terms of prediction accuracy. Despite research also tries to a variety of machine learning methods applied to landslide susceptibility cartography, and mainly depends on the comparison of the quantitative indicators evaluation, although these measures have certain reference significance, however, is difficult to fully reflect the effectiveness and reliability of the model. In contrast, the integrated mapping method proposed in this study effectively reduces the uncertainty caused by relying on a single model by combining the advantages of different models.

5 Discussion

Different modeling methods often produce inconsistent spatial distribution results in landslide susceptibility prediction, which makes it difficult to optimize the prediction map in disaster risk management. In order to alleviate this problem, this study proposes a method to integrate effective information by quantifying the spatial consistency between models and fusing multiple prediction results to improve the reliability of landslide prone zoning. Taking Ruijin City, a high landslide incidence area in Jiangxi Province, as a case study, three machine learning methods including SVM, XGBOOST and Random Forest (RF) were used to generate the landslide susceptibility distribution map based on hyperparameter tuning. Compared with previous studies on this area, this study used the updated impact factor data, and excluded low-lying areas (including water bodies and areas below 5 m above sea level) in the mapping process to avoid overestimation of the landslide prone range as in some recent studies. The performance evaluation showed that all the models showed excellent predictive ability, with AUC values ranging from 0.84 to 0.93, which was consistent with the conclusions of multiple current studies on the high accuracy of machine learning in landslide prediction. This study suggests that, even based on the same set of cataloged landslide data, different modeling methods may still generate spatially diverse prediction results. Based on the analysis of the spatial consistency of the model outputs, it was found that there was obvious regional heterogeneity in landslide prediction, and the pixel-by-pixel correlation coefficients between the models ranged from 0.78 to 0.84. In addition, there are some differences between the prone zone map drawn in this study and another recent study on Ruijin city. These differences indicate that although each model shows high AUC values and excellent classification performance, there are still uncertainties in the prediction results that cannot be ignored. Most of the current researches on landslide susceptibility based on machine learning focus on the optimal prediction method, but pay less attention to the uncertainty caused by the spatial inconsistency between different models. Based on an actual case, this study is the first to systematically investigate this issue, and fills the gap of existing research in related fields.

By integrating the prediction results of multiple machine learning landslide susceptibility models, this study aims to reduce the uncertainty in the prediction and improve the accuracy of landslide spatial probability assessment. The proposed modeling framework is transferable and can be applied to other landslide prone areas. Landslide susceptibility maps (LSM) can provide a scientific basis for urban planners to identify suitable areas for construction. Taking Ruijin City as an example, the comprehensive susceptibility map generated in this study can assist policy makers and engineers to determine the implementation focus and timing of landslide risk management measures. The model is an improvement of the existing landslide prediction methods, which is helpful to achieve more accurate spatial prediction of landslide hazard. The output of the model can be used to optimize the landslide warning system, thereby enhancing the effectiveness of disaster risk mitigation strategies and supporting local communities to build a more disaster resilient development environment.

Statements

Data availability statement

The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.

Author contributions

XZ: Data curation, Writing – original draft. YX: Formal Analysis, Funding acquisition, Writing – review and editing.

Funding

The author(s) declare that financial support was received for the research and/or publication of this article. This research was funded by the CRSRI Open Research Program (Program SN:CKWV20241187/KY); Key Laboratory of Land Satellite Remote Sensing Application, Ministry of Natural Resources of the People’s Republic of China (KLSMNR-G202304); Key Laboratory of Coastal Salt Marsh Ecosystems and Resources, Ministry of Natural Resources (KLCSMERMNR202306); Suzhou University of Science and Technology Talent Introduction Initiation Project (332214808); Vice President of Science and Technology of Jiangsu Province (1106); The Natural Science Foundation of the Jiangsu Higher Education Institutions of China (25KJB420006).

Acknowledgments

The authors would like to thank the Department of Surveying and Mapping of Jiangxi Province for providing relevant data.

Conflict of interest

Author XZ was employed by Nanjing Changtian Surveying and Mapping Technology Co., Ltd.

The remaining author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declare that no Generative AI was used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

  • 1

    AhmadS.RizviZ. H.ArpJ. C. C.WuttkeF.TirthV.IslamS. (2021). Evolution of temperature field around underground power cable for static and cyclic heating. Energies14 (23), 8191. 10.3390/en14238191

  • 2

    AhmadH.YaqubM.LeeS. H. (2024). Environmental-social-and governance-related factors for business investment and sustainability: a scientometric review of global trends. Environ. Dev. Sustain.26 (2), 29652987. 10.1007/s10668-023-02921-x

  • 3

    AhmadS.AhmadS.AkhtarS.AhmadF.AnsariM. A. (2025a). Data-driven assessment of corrosion in reinforced concrete structures embedded in clay dominated soils. Sci. Rep.15, 22744. 10.1038/s41598-025-08526-w

  • 4

    AhmadS.PrakashR.AbediS. (2025b). “Microscale chemo-mechanics of shale-reactive brine interaction,” in The Biot-Bazant Conference. 2021, Conference contribution. 10.6084/m9.figshare.14787516.v1

  • 5

    AhmadS.RizviZ. H.WuttkeF. (2025c). Unveiling soil thermal behavior under ultra-high voltage power cable operations. Sci. Rep.15, 7315. 10.1038/s41598-025-91831-1

  • 6

    AhmedB. (2021). The root causes of landslide vulnerability in Bangladesh. Landslides18 (5), 17071720. 10.1007/s10346-020-01606-0

  • 7

    AlamE.Ray-BennettN. S. (2021). Disaster risk governance for district-level landslide risk management in Bangladesh. Int. J. Disaster Risk Reduct.59, 102220. 10.1016/j.ijdrr.2021.102220

  • 8

    Alcántara-AyalaI. (2025). Landslides in a changing world. Landslides22, 28512865.

  • 9

    BaoH.ZengC.PengY.WuS. (2022). The use of digital technologies for landslide disaster risk research and disaster risk management: progress and prospects. Environ. Earth Sci.81 (18), 446. 10.1007/s12665-022-10575-7

  • 10

    BatarA. K.WatanabeT. (2021). Landslide susceptibility mapping and assessment using geospatial platforms and weights of evidence (WoE) method in the Indian Himalayan region: recent developments, gaps, and future directions. ISPRS Int. J. Geo-Information10 (3), 114. 10.3390/ijgi10030114

  • 11

    CuiS.WuH.PeiX.YangQ.HuangR.GuoB. (2022). Characterizing the spatial distribution, frequency, geomorphological and geological controls on landslides triggered by the 1933 Mw 7.3 Diexi earthquake, Sichuan, China. Geomorphology403, 108177. 10.1016/j.geomorph.2022.108177

  • 12

    DahmaniL.LaaribyaS.NaimH.TunguzV.DindarogluT. (2024). Assessing landslide susceptibility in Chefchaouen, Morocco: an application of the landslide numerical risk factor method for sustainable urban development and disaster risk management. Biosyst. Divers.32 (3), 389397. 10.15421/012442

  • 13

    GuoF. Y.MengX. M.QiT. J.DijkstraT.ThorkildsenJ. K.YueD. x.et al (2022). Rapid onset hazards, fault-controlled landslides and multi-method emergency decision-making. J. Mt. Sci.19 (5), 13571369. 10.1007/s11629-021-6941-x

  • 14

    JaafariA. (2024). Landslide susceptibility assessment using novel hybridized methods based on the support vector regression. Ecol. Eng.208, 107372. 10.1016/j.ecoleng.2024.107372

  • 15

    KabA.DjerbalL.BaharR. (2023). Implementation of PCA multicollinearity method to landslide susceptibility assessment: the study case of Kabylia region. Arabian J. Geosciences16 (4), 291. 10.1007/s12517-023-11374-5

  • 16

    KavzogluT.TekeA. (2022a). Predictive performances of ensemble machine learning algorithms in landslide susceptibility mapping using random forest, extreme gradient boosting (XGBoost) and natural gradient boosting (NGBoost). Arabian J. Sci. Eng.47 (6), 73677385. 10.1007/s13369-022-06560-8

  • 17

    KavzogluT.TekeA. (2022b). Advanced hyperparameter optimization for improved spatial prediction of shallow landslides using extreme gradient boosting (XGBoost). Bull. Eng. Geol. Environ.81 (5), 201. 10.1007/s10064-022-02708-w

  • 18

    KumarC.WaltonG.SantiP.LuzaC. (2023). An ensemble approach of feature selection and machine learning models for regional landslide susceptibility mapping in the arid mountainous terrain of southern Peru. Remote Sens.15 (5), 1376. 10.3390/rs15051376

  • 19

    LiH.SamsudinN. A. (2024). A systematic review of landslide research in urban planning worldwide. Nat. Hazards121, 63916411. 10.1007/s11069-024-07064-4

  • 20

    LinN.ZhangD.FengS.DingK.TanL.WangB.et al (2023). Rapid landslide extraction from high-resolution remote sensing images using SHAP-OPT-XGBoost. Remote Sens.15 (15), 3901. 10.3390/rs15153901

  • 21

    LiuR.YangX.XuC.WeiL.ZengX. (2022). Comparative study of convolutional neural network and conventional machine learning methods for landslide susceptibility mapping. Remote Sens.14 (2), 321. 10.3390/rs14020321

  • 22

    LiuQ.TangA.HuangD. (2023). Exploring the uncertainty of landslide susceptibility assessment caused by the number of non–landslides. Catena227, 107109. 10.1016/j.catena.2023.107109

  • 23

    LuZ.LiuG.SongZ.SunK.LiM.ChenY.et al (2024). Advancements in technologies and methodologies of machine learning in landslide susceptibility research: current trends and future directions. Appl. Sci.14 (21), 9639. 10.3390/app14219639

  • 24

    Morales-HernándezA.Van NieuwenhuyseI.Rojas GonzalezS. (2023). A survey on multi-objective hyperparameter optimization algorithms for machine learning. Artif. Intell. Rev.56 (8), 80438093. 10.1007/s10462-022-10359-2

  • 25

    Pacheco QuevedoR.Velastegui-MontoyaA.Montalván-BurbanoN.Morante-CarballoF.KorupO.Daleles RennóC. (2023). Land use and land cover as a conditioning factor in landslide susceptibility: a literature review. Landslides20 (5), 967982. 10.1007/s10346-022-02020-4

  • 26

    SahaS.RoyJ.PradhanB.HembramT. K. (2021). Hybrid ensemble machine learning approaches for landslide susceptibility mapping using different sampling ratios at east Sikkim himalayan, India. Adv. Space Res.68 (7), 28192840. 10.1016/j.asr.2021.05.018

  • 27

    SajidM.Płotka-WasylkaJ. (2022). Green analytical chemistry metrics: a review. Talanta238, 123046. 10.1016/j.talanta.2021.123046

  • 28

    SelamatS. N.MajidN. A.TahaM. R. (2025). Multicollinearity and spatial correlation analysis of landslide conditioning factors in langat River basin, Selangor. Nat. Hazards121 (3), 26652684. 10.1007/s11069-024-06903-8

  • 29

    SharmaN.SahariaM.RamanaG. V. (2024). High resolution landslide susceptibility mapping using ensemble machine learning and geospatial big data. Catena235, 107653. 10.1016/j.catena.2023.107653

  • 30

    ShuX.YeY. (2023). Knowledge discovery: methods from data mining and machine learning. Soc. Sci. Res.110, 102817. 10.1016/j.ssresearch.2022.102817

  • 31

    SousaJ. J.LiuG.FanJ.PerskiZ.StegerS.BaiS.et al (2021). Geohazards monitoring and assessment using multi-source earth observation techniques. Remote Sens.13 (21), 4269. 10.3390/rs13214269

  • 32

    StukeA.RinkeP.TodorovićM. (2021). Efficient hyperparameter tuning for kernel ridge regression with Bayesian optimization. Mach. Learn. Sci. Technol.2 (3), 035022. 10.1088/2632-2153/abee59

  • 33

    TehraniF. S.CalvelloM.LiuZ.ZhangL.LacasseS. (2022). Machine learning and landslide studies: recent advances and applications. Nat. Hazards114 (2), 11971245. 10.1007/s11069-022-05423-7

  • 34

    TengF.MaoY.LiY.QianS.NanehkaranY. A. (2024). Comparative models of support-vector machine, multilayer perceptron, and decision tree predication approaches for landslide susceptibility analysis. Open Geosci.16 (1), 20220642. 10.1515/geo-2022-0642

  • 35

    UllahI.AslamB.ShahS. H. I. A.TariqA.QinS.MajeedM.et al (2022). An integrated approach of machine learning, remote sensing, and GIS data for the landslide susceptibility mapping. Land11 (8), 1265. 10.3390/land11081265

  • 36

    WangY.NanehkaranY. A. (2024). GIS-based fuzzy logic technique for mapping landslide susceptibility analyzing in a coastal soft rock zone. Nat. Hazards120 (12), 1088910921. 10.1007/s11069-024-06649-3

  • 37

    WangX.ZhangC.WangC.LiuG.WangH. (2021). GIS-based for prediction and prevention of environmental geological disaster susceptibility: from a perspective of sustainable development. Ecotoxicol. Environ. Saf.226, 112881. 10.1016/j.ecoenv.2021.112881

  • 38

    WangS.ZhuangJ.ZhengJ.FanH.KongJ.ZhanJ. (2021). Application of Bayesian hyperparameter optimized random forest and XGBoost model for landslide susceptibility mapping. Front. Earth Sci.9, 712240. 10.3389/feart.2021.712240

  • 39

    WangT.ZilinskasR.LiY.QuY. (2023). Missing data imputation for a multivariate outcome of mixed variable types. Statistics Biopharm. Res.15 (4), 826837. 10.1080/19466315.2023.2169753

  • 40

    WeiA.YuK.DaiF.GuF.ZhangW.LiuY. (2022). Application of tree-based ensemble models to landslide susceptibility mapping: a comparative study. Sustainability14 (10), 6330. 10.3390/su14106330

  • 41

    XuZ.CheA.ZhouH. (2024). Seismic landslide susceptibility assessment using principal component analysis and support vector machine. Sci. Rep.14 (1), 3734. 10.1038/s41598-023-48196-0

  • 42

    ZengL.Brignardello-PetersenR.HultcrantzM.SiemieniukR. A.SantessoN.TraversyG.et al (2021). GRADE guidelines 32: GRADE offers guidance on choosing targets of GRADE certainty of evidence ratings. J. Clin. Epidemiol.137, 163175. 10.1016/j.jclinepi.2021.03.026

  • 43

    ZhaiS.DuG.HeH. (2024). Bayesian probabilistic characterization of the shear-wave velocity combining the cone penetration test and standard penetration test. Stoch. Environ. Res. Risk Assess.38 (1), 6984. 10.1007/s00477-023-02566-2

Summary

Keywords

uncertainty, landslide, prediction, susceptibility, machine learning (ML)

Citation

Zhou X and Xing Y (2025) Spatial consistency assessment and landslide susceptibility prediction optimization. Front. Earth Sci. 13:1702688. doi: 10.3389/feart.2025.1702688

Received

10 September 2025

Accepted

17 October 2025

Published

31 October 2025

Volume

13 - 2025

Edited by

Hans-Balder Havenith, University of Liège, Belgium

Reviewed by

Zarghaam Rizvi, GeoAnalysis Engineering GmbH, Germany

Sudesh Pundir, Pondicherry University, India

Updates

Copyright

*Correspondence: Yin Xing,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics