ORIGINAL RESEARCH article

Front. Built Environ., 12 May 2026

Sec. Structural Engineering and Design

Volume 12 - 2026 | https://doi.org/10.3389/fbuil.2026.1811594

Comparative evaluation of supervised machine learning models in the non-destructive strength prediction of steel fiber reinforced concrete

  • 1. Department of Civil Engineering, National Institute of Technology, Calicut, India

  • 2. Department of Civil Engineering, Christ University, Bangalore, India

Abstract

This study evaluates the performance of nine supervised machine learning (ML) techniques, namely, K-nearest neighbors (KNN), decision tree (DT), support vector regression (SVR), random forest (RF), gradient boosting (GB), AdaBoost (AB), extreme gradient boosting (XGBoost), light gradient boosting machine (LGBM), and categorical boosting (CatBoost)—on predicting the compressive and tensile strengths of steel fiber reinforced concrete (SFRC), thus providing a data-driven, non-destructive framework for assessing material performance. An extensive dataset is created by varying the fine-to-coarse aggregate ratio and steel fiber content in the range of 0%–2%. The models were evaluated using performance metrics such as coefficient of determination, R2, and mean squared error (MSE). The results indicate that XGBoost exhibited the best performance of all the compared ML models, as evident through its R2 value of 0.9926 for flexural strength, 0.9965 for split tensile, and 0.7837 for compressive strength prediction. These results illustrate that the nonlinear behavior of SFRC can be captured more effectively by AI-ML models and provide better and more accurate strength predictions, thereby supporting advanced non-destructive testing strategies and reducing reliance on extensive destructive testing. Through this study, ML models can be positioned within a structural health monitoring (SHM) context where the predicted parameters of strength can be used for maintenance planning, condition assessment, and damage detection. The outcomes of this study contribute to the enhancement of data-driven approaches for material characterization, thus helping incorporate ML models into real-world structural assessment frameworks.

1 Introduction

Steel fiber reinforced concrete (SFRC) is a high-performance concrete with short, discrete lengths of steel fiber with different aspect ratios. It has extensive applications in civil engineering, especially in cases where there is a need for increased strength, crack resistance, and ductility. SFRC is widely used in construction such as pavements, bridges, industrial floors, and precast. The property of discontinuity increases the tensile and flexural strengths of SFRC, and its high crack resistance has made it a leading high performance construction material. The presence of discontinuous fibers adds significantly to the performance level of SFRC compared to normal concrete. The nonlinear behavior of SFRC due to the presence of fibers, varying aspect ratio, matrix composition, and curing conditions makes its mechanical behavior very complex. Because of this complexity, it becomes very challenging to predict its mechanical properties using conventional empirical or egression approaches. Due to their ability to capture nonlinear relationships among input parameters, machine learning (ML) models have gained significant attention for predicting the material properties of concrete in recent years. Algorithms such as artificial neural networks (ANNs), support vector machines (SVMs), random forest (RF), and gradient boosting techniques have been used successfully for mechanical strength prediction, including the tensile, flexural, and compressive strength of concrete. However, most studies are limited as they focus is on a small subset of ML models, or dataset heterogeneity is observed with changing preprocessing techniques. Such discrepancies hamper a fair comparison and do not allow for a generic reproducibility of results. Even though there is a considerable literature on the potential of ML models within a SHM context, studies lack practical integration, with more emphasis being laid on conceptual aspects. The use of ML-based strength prediction for structural health assessment, damage detection, and maintenance planning has not yet been explored. To bridge this gap, there is a need for a systematic evaluation framework, which not only predicts strength but also has a clear correlation between SHM objectives and predicted material properties. Another important gap that has been identified in the literature is the absence of a standardized dataset, which has multiple strength predictors applicable to engineering practice. For reliability-based design and assessment, it is necessary to incorporate statistical measures such as minimum, maximum, and characteristic strengths, which are lacking in the majority of studies since they focus more on single output predictions such as compressive strength. To overcome these gaps, the present study focuses on providing a systematic and application-oriented evaluation of strength prediction using ML-based models for SFRC. A standard dataset comprising 432 samples has been developed by integrating critical material parameters and various strength descriptors. Using a unique preprocessing and validation framework, nine ML algorithms have been evaluated, enabling a fair comparison and benchmarking of its predictive performance. The models are also evaluated using statistical metrics to guarantee a robust and fair comparison. This study also focuses on positioning the developed models within a SHM context not limited to conventional prediction alone. The integration of the predicted values with non-destructive testing (NDT) data can act as proxy indicators to support structural condition assessment, damage detection, and maintenance planning. This permits a practical framework that integrates data-driven models in structural engineering applications. An extensive literature review is presented in Section 2 on the different types of concrete and their application areas, SHM concepts within the non-destructive testing context, and the ML models. The methodology we adopt, along with the objectives and material testing, is explained in Section 3. Section 4 presents a detailed analysis and comparative evaluation of the nine ML models supported by results, with conclusions being highlighted in Section 6 and future scope of this study in Section 7. A detailed abbreviation table is provided in Section 8 for ease of the reader.

2 Literature review

Evaluation of the performance of civil infrastructure along with continuous structural health monitoring requires the accurate assessment of the mechanical properties of concrete. The experimental methods used for assessing these properties are often found to be time-consuming, resource-intensive, and destructive in nature. SFRC is being extensively adopted, which involves the interaction of multiple mix design parameters, thus leading to even more complicated property evaluation mechanisms. In this review, the tensile, flexural, and compressive strengths of SFRC are predicted using ML techniques. The different techniques used here are K-nearest neighbors (KNN), decision tree (DT), support vector regression (SVR), random forest (RF), gradient boosting (GB), AdaBoost (AB), extreme gradient boosting (XGBoost), light gradient boosting machine (LGBM), and categorical boosting (CatBoost).

2.1 Engineered concrete

Trends in concrete technology which attempt to increase its mechanical properties have given rise to different types of concrete, such as high-strength and high-performance concrete (; ; ; ). Concrete with compressive strength greater than 60 MPa and up to about 150 MPa is generally classified as high-strength. This is achieved by reducing the water–cement (W/C) ratio below 0.35. However, lowering the W/C ratio reduces the concrete’s workability. This is overcome with the addition of superplasticizers to obtain the desired workability.

The advantages of high-strength concrete allow the resisting of higher loads for optimized sections of beams, columns, slabs, and floors. This in turn reduces the cost of formwork (). A concrete mixture with high strength, durability, workability, modulus, and dimensional stability, low permeability, and resistance to chemical attack is generally said to be “high-performance”. Though it has a higher initial cost than conventional concrete, its maintenance cost and durability add to its cost benefits (). It is used for several construction purposes, especially for severe exposure conditions. The materials used include silica fume, granulated blast furnace slag, fly ash, and steel fiber (Yazıcı et al., 2007).

2.2 Structural health monitoring and non-destructive testing

One of the most essential components of identifying potential failure in civil structures is structural health monitoring (SHM) (). This also encompasses the management of modern infrastructure, thereby ensuring the early detection of failures and damage. Modern structural health monitoring systems consist of non-destructive testing (NDT) techniques, which enable the assessment of material properties without compromising structural integrity (). This is in contrast with destructive testing methods, which extract specimens, thus resulting in local damage. NDT methods offer a fast, cost-effective, and scalable solution for evaluating structural performance. The common degradation mechanisms that occur in concrete structures, such as cracking, void formation, and material deterioration, can be monitored using NDT techniques in addition to the detection of internal defects and the estimation of key mechanical properties. Ultrasonic pulse velocity (UPV), rebound hammer testing, impact-echo, acoustic emission, and resonant frequency measurements are some of the common NDT techniques widely used for monitoring concrete structures wherein periodic or continuous monitoring is essential over time (). The ability of NDT techniques to detect damage-sensitive parameters make them more suitable for performance-based decision making, planning maintenance activities, and condition assessment ().

The major challenge in utilizing NDT measurements for evaluating mechanical properties is to establish a relationship between NDT parameters and mechanical properties such as tensile, flexural, and compressive strength. This is attributed to the fact that NDT measurements are often influenced by external and environmental parameters like the moisture content, characteristics of the aggregates used, distribution of the fiber content, conditions used for curing, and age of the concrete. When SFRC is being used, this challenge increases because of the additional heterogeneity and nonlinear behavior introduced by the presence of discrete fibers in SFRC (Su et al., 2024). SFRC structures that are exposed to aggressive environments or those which are subjected to dynamic loading can adopt NDT techniques for periodic monitoring. The volume of the fiber used, its aspect ratio, the characteristics of the bonding formed, and their orientation all impact the measurements obtained from NDT techniques, especially those related to wave propagation and surface hardness. As a result, traditional correlation models are unable to effectively capture the interactions between NDT parameters and the mechanical response of SFRC, thereby leading to the minimal use of SHM applications involving SFRC (). The emergence of ML models has led to their adoption in data-driven approaches to more meaningfully interpret the data obtained from NDT (Wang et al., 2025; Wijesundara et al., 2025; Zheng et al., 2022; Vihas et al., 2025). The development of experimental datasets has paved the way for learning complex nonlinear relationships, and thus these approaches offer a significant increase in predictive capability and adaptability compared to traditional methods.

2.3 Artificial intelligence-based regression models

The emergence of machine learning (ML) and deep learning artificial intelligence techniques have led to breakthroughs in civil engineering as in other areas. Research on the compressive strength of concrete with various algorithms supports this. The difference is that the concrete under investigation varies based on composition. The artificial neural network (ANN) methodology with the Levenberg Marquardt algorithm was considered to be an accurate method for predicting compressive strength as reported by and by for a self-cured concrete. identified ANN with the Bayesian regularization algorithm as the most accurate for high-performance concrete.

Through Monte Carlo simulation and statistical analysis, observed that GPR-32 (Gaussian Process Regression kernel 32) is more efficient than other algorithms for predicting the compressive strength of high-performance concrete. Several studies have been undertaken on the compressive strength of concrete and high-performance concrete, but few have been done on split tensile and flexural strength. Yan et al. (2013) identified support vector machines (SVMs) for predicting the split tensile strength of concrete from its compressive strength. found XGBoost and gradient boost regressors to be the most appropriate ML algorithms for predicting both the flexural and compressive strengths of SFRC. attempted to predict all three strengths of an SFRC using ANN and observed that the Genetic Algorithm gave better results in terms of overall predictive efficiency but was outperformed by incremental back propagation (IBP) and batch back propagation (BBP) for predicting flexural and compressive strengths. Thus, the need to develop an artificial intelligence model for predicting the three different strengths of SFRC is evident from the latest literature.

In this study, an attempt was made to understand the various mechanical properties of SFRC by varying their fiber content without the addition of any other supplementary cementitious material apart from the steel fiber. The effects of the various parameters on the properties of SFRC are considered and analyzed with the help of artificial intelligence methods. The study also seeks to shed light on the algorithm used to present a better artificial intelligence model for predicting the mechanical properties of concrete.

The KNN algorithm is a supervised ML algorithm which learns a function from a set of labelled input data, thereby producing an appropriate output using non-labelled data. It is most commonly used to classify problems and also for regression (). Decision tree is also a supervised ML algorithm used for both classification and regression problems. Its functioning is similar to a tree, wherein it starts with the root node (dataset), which expands on further branches (rules) and constructs a tree-like structure with leaves (outcome) (). Support vector regression is a supervised ML algorithm based on SVMs. It is used for classification as well as regression problems, primarily for classification problems in machine learning. The SVM algorithm creates the best line or decision boundary, called a “hyperplane,” that segregates n-dimensional space into classes to handle or classify the new data point in the correct category in the future (). Random forest also falls into the category of a supervised machine learning algorithm. It builds an ensemble of decision trees by training a combination of learning models to increase the overall result (). The gradient boosting algorithm () is a sequential ensemble learning algorithm where the performance of the model enhances over iterations. In ML, this algorithm is used to solve classification and regression problems, combining multiple weak models to achieve better performance. It is a highly robust technique and is employed in several risk functions to optimize the accuracy of the model’s prediction.

Yennimar et al. (2024) highlight that AdaBoost, or adaptive boosting, is the first boosting ensemble algorithm (building a strong model from multiple weak models) and is mainly used for binary classification. Extreme gradient boosting, or the XGBoost, algorithm () is a gradient-boosted decision tree machine learning algorithm. It is designed to enhance the performance and speed of a ML model. XGBoost is an ensembled supervised machine learning algorithm with an extended version of gradient boosting and decision trees. Light gradient boosting machine, or LightGBM, is also a gradient-boosted decision tree ML algorithm. It is named thus due to its faster prediction of outputs and computation of power (). The categorical boosting, or CatBoost, algorithm is also a gradient-boosted decision tree ML algorithm (). It uses symmetric trees to provide faster execution. CatBoost is a high-performance gradient boosted algorithm with improved accuracy and faster predictions.

The importance of using multiple performance metrics along with systematic model comparison and validation have been emphasized in recent research on data-driven modeling. highlight the improvement in the robustness and generalizability of predictions by evaluating a large number of ML models under consistent conditions. They stress the use of diverse parameters as input and adopting comprehensive evaluation metrics to evaluate the model’s performance.

A comparison of recent ML-based studies on predicting concrete strength is presented in Table 1, highlighting the methodologies used, the dataset size, and some key limitations. This helps us to arrive at the research gap is presented in Section 2.4.

TABLE 1

StudyMaterial typeML models usedDataset sizeOutput parameterSHM integrationLimitations
Zhang et al. (2025)Self-compacting concretePCA, RVM, PSONot specifiedCompressive strengthNoNo comparative study
Vihas et al. (2025)Concrete with GGBSSVM, RF, ANN, XGBoost560Compressive strengthNoOnly compressive strength is predicted
ConcreteExplainable ML modelsNot specifiedCompressive strengthNoFocus is on interpretability alone
Geopolymer concreteANN, XGBoost, LSTMNot specifiedCompressive strengthNoFocus on specific material type
High-performance concreteRF, GRNN300Compressive strengthNoLimited model diversity for benchmarking
Yang et al. (2024)Fly ash concreteANN, SVM, XGBoost200Compressive strengthNoLimited dataset and single parameter focus
ConcreteRF, SVM, XGBoost, ANN1030Compressive strengthNoOnly compressive strength prediction
ConcreteRF, ANN, XGBoost, etc.400Compressive strengthLimited (GUI-based)Focus on platform development

Summary of existing studies on machine learning for concrete strength prediction.

2.4 Research gaps identified

From the above literature review, it is clear that existing studies focus on the prediction of compressive strength alone and also that the ML models used are limited. The comparison is unfair because of the use of heterogeneous datasets and inconsistent validation approaches. Furthermore, it is evident from Table 1 that a SHM framework with ML integration is yet in the conceptual stage. These gaps highlight the need for a systematic, reproducible, and application-oriented comparative study. Also motivated by the methodology adopted in , this study presents a structured and unified benchmarking of nine ML algorithms for which a curated dataset has been developed. The predicted output parameters are linked to a SHM-based assessment of conditions which addresses the gaps identified in the literature.

3 Materials and methods

Although several studies have been undertaken regarding the various properties of high-performance concrete, there is still a vast uninvestigated area in concrete technology. The greatest challenge is the lack of a large body of data and the inability to choose the right ML or deep learning models. Most studies are limited to a certain number of algorithms and fail to investigate the possibility of better algorithms. From the literature review, it is evident that the studies concerning the flexural and tensile strengths of high-performance concrete using artificial intelligence models are very limited. The studies that deal with the fresh and hardened properties of SFRC also differ from each other due to the variation in the supplementary cementitious material used, such as silica fume, and glass-burnt furnace slag. Therefore, a dataset is needed that purely contains SFRC in addition to normal concrete mixture.

3.1 Objectives of the study

In alignment with the NDT-based SHM framework, the objectives for the present study are formulated as follows.

  • Development of NDT-based dataset for machine learning. To develop a comprehensive dataset incorporating non-destructive testing (NDT) parameters along with corresponding split tensile strength, flexural strength, and compressive strength of steel fiber reinforced concrete (SFRC).

  • Analysis of key NDT and material input parameters, including fine-to-coarse aggregate ratio and steel fiber content, on the mechanical properties of SFRC, with particular emphasis on structural condition assessment.

  • Development and validation of machine learning (ML) models for SHM applications. To develop, implement, and validate artificial intelligence-based predictive models using Python programming for NDT-driven estimation of SFRC mechanical properties, and to identify the most suitable ML algorithm for structural health monitoring (SHM) applications through verification against experimental data.

3.2 Experimental database

To predict the mechanical properties of SFRC within a NDT-based SHM framework, we developed an experimental database. The dataset consists of experimentally tested SFRC specimens with recorded mix-design parameters, NDT measurements, and corresponding mechanical properties. The first phase consisted of designing the mix proportion of M25, M30, M35, and M40 grades of concrete. The concrete specimens were cast based on the mix proportion designed by varying the percentage of steel fiber content from 0%, 1.0%, 1.5 % to 2.0% and also the fine to coarse aggregate ratio varying between 0.6, 0.5, and 0.7 (for all grades except M40—0.5, 0.4, 0.6) for each grade of concrete. We also determined properties such as density, slump, vee-bee, and the strength properties of the specimens for a maximum of 28 days.

3.3 Dataset description

The dataset curated for the this study has the parameters as tabulated in Table 2.

TABLE 2

ParameterValue
Total number of samples432
Training data80%
Testing data20%
Validationk-fold CV

Dataset and model configuration parameters.

3.4 Material testing

Split tensile strength was determined using the indirect tensile test method (IS 5816) (), while flexural strength was evaluated through third-point loading tests on prism specimens (IS 516). Compressive strength tests were conducted on standard specimens in accordance with established testing procedures (IS 516). All tests were performed at the specified curing age under controlled laboratory conditions.

3.5 Specific gravity tests

As per the appropriate IS codes, specific gravity tests were conducted to determine values for cement, fine aggregate, and coarse aggregate. The specific gravity of cement was evaluated using Le Chatelier’s apparatus (); the value obtained was 3.15. Similarly, specific gravity tests were conducted on fine aggregates and coarse aggregates; the values obtained were 2.52 and 2.6, respectively. All values obtained were well within the permissible limits. Tables 35 provide the values obtained for these tests with the three trials conducted. The physical properties of the steel fiber used in SFRC are given in Table 6. The mix proportions of SFRC with the ratio used is as provided in Table 7.

TABLE 3

SpecificationTrial 1Trial 2Trial 3
Weight of empty flask (g)121121122
Weight of flask + cement (g)173172171.5
Weight of flask + cement + kerosene (g)382382381.5
Weight of flask + kerosene (g)343344344.5
Weight of water + flask (g)401400401
Specific gravity of kerosene0.7930.7990.797
Specific gravity of cement3.173.133.16

Specific gravity of cement.

TABLE 4

SpecificationTrial 1Trial 2Trial 3
Weight of sample (g)500500500
Weight of pycnometer + sample + water (g)178017951800
Weight of pycnometer + water (g)148714871487
Weight of oven-dry sample (g)493490490
Specific gravity2.382.552.62
Water absorption (%)1.422

Specific gravity and water absorption of aggregate.

TABLE 5

SpecificationTrial 1Trial 2Trial 3
Weight of the saturated sample suspended in wire basket (g)2.0852.0652.02
Weight of the basket suspended in water (g)0.830.830.815
Weight of surface dry aggregate in air (g)221.99
Dry weight of aggregate (g)21.991.99
Weight of saturating aggregate in water (g)1.2551.2351.205
Weight of water equal to volume of aggregate (g)0.7450.7650.785
Specific gravity of aggregate2.72.62.5
Water absorption of aggregate (%)00.50

Specific gravity and water absorption of coarse aggregate.

TABLE 6

SpecificationsTypeDia (mm)UTS (kg/mm2)Length (mm)
Carbon steel fibreHooked end0.75128.2150

Properties of carbon steel fiber.

TABLE 7

MaterialQuantity (kg/m3)
Cement372 kg/m3
Fine aggregate657 kg/m3
Coarse aggregate1106 kg/m3
Water186 kg/m3
Water–cement ratio0.5
Superplasticizer0.3% (by weight of cement)
Steel fiber1.5% and 2% for fine-to-coarse aggregate ratio of 0.4

Mix proportions of SFRC.

3.6 Mix proportion design

The mix proportion was prepared as per IS 456:2000 () and IS 10262:2019 () for four different grades of concrete: M25, M30, M35, and M40. Table 5 below shows the design data considered for M30 grade concrete.

The mix proportion for each of the M25, M30, M35, and M40 grade concretes was designed in accordance with the IS 10262 and IS 456 codes. Using this mix proportion, it was possible to achieve expected characteristic compressive strength, durability, and workability. The proportions of cement, fine and coarse aggregates, and water were optimized based on the target mean strength and selected water–cement ratio. The final mix proportion that was obtained was in line with the workability and strength requirement for each grade of concrete, thus warranting the effective performance and consistency expected in structural applications. The concrete specimens were cast as per IS516 (). The size of the cubes was 150 mm × 150 mm × 150 mm, and the number of samples used was three cubes for 7 days of testing, three cubes for 14 days of testing, and three cubes for 28 days of testing. The size of the cylinder was 150 mm in diameter and 300 mm in length, as per IS 5816 (1999): Method of Test Splitting Tensile Strength of Concrete (). The number of samples used was three cylinders for 28 days of testing. The size of the beams was 700 mm in length, and the cross-section was 150 mm × 150 mm, as per IS 516, and three sample beams were used for 28 days of testing.

3.7 Hyperparameter tuning

To ensure optimal model performance and avoid overfitting, hyperparameter tuning was performed using the Grid Search approach with the K-Fold cross validation method. To ensure fair comparison among all models, the same tuning strategy and evaluation criteria were applied. The computational experiments were conducted using the same hardware: Intel i7 processor and 16 GB RAM with Google Colab. Table 8 below shows the hyperparameter tuning details for all ML models used in the study.

TABLE 8

ModelHyperparameterValue considered
Support vector machine (SVM)KernelLinear
Regularization10
Gamma0.01
Artificial neural networks (ANNs)Number of hidden layers2
Neurons per layer50
Learning rate0.001
Epochs150
Batch size32
Random forest (RF)Number of trees200
Maximum depth10
Minimum samples split2
Minimum samples leaf1
XGBoostNumber of estimators300
Learning rate0.01
Maximum depth10
Subsample0.8
LightGBMNumber of leaves50
Learning rate0.01
Maximum depth10
Feature fraction0.8
CatBoostIterations300
Learning rate0.01
Depth10
k-nearest neighbors (KNN)Number of neighbors5
Distance metricEuclidean
Gradient boosting (GBM)Number of estimators200
Learning rate0.01
Maximum depth10
Decision tree (DT)Maximum depth10
Minimum samples split2

Hyperparameter settings of machine learning models.

3.8 Machine learning modeling

ML modeling was developed using the Python programming language in Google Colab notebook. The nine algorithms were selected based on the previous literature reviews and were applied for the accuracy test. The training and the testing accuracy were evaluated using the mean squared error (MSE) prediction error tests (Equation 1) and coefficient of determination (Equation 2):where MSE is the mean squared error, n is the number of data points, Yi is the observed values, and Ŷi represents the predicted values;where is the coefficient of determination, represents the observed values, represents the predicted values, and is the mean value of Y.

4 Results and discussion

Flexural testing of the concrete samples was conducted using the Universal Testing Machine with a maximum capacity of 1000 kN. A total of 432 samples of concrete cubes were cast for different variations of grade of concrete, steel fiber content, and fine-to-coarse aggregate ratio for determining the compressive strength of concrete. Similarly, 144 samples each were cast for determining the split tensile strength and flexural strength of the concrete. The compressive and split tensile tests of the concrete samples were conducted using the Compressive Testing Machine with a maximum capacity of 2000 kN.

For a total of 432 samples, the various characteristics of the dataset were then tabulated. Table 9 shows the mean, standard, minimum, and maximum values of the input (cement, water, fine aggregate, coarse aggregate, superplasticizer, steel fiber, density, age, slump, and vee bee) and output (flexural, split tensile, and compressive strengths) variables for the experimental dataset obtained.

TABLE 9

ParticularCountMeanStdMinMax
Cement432421.7555.51375.00510.00
Water432203.922.40200.00206.00
Coarse aggregate432664.7365.18576.00772.80
Fine aggregate4321173.50126.401073.001440.00
Superplasticizer4320.440.450.001.53
Steel fiber4324.903.300.0010.20
Density4322292.1937.062168.352382.22
Age43216.338.747.0028.00
Slump43281.9311.2858.00115.00
Vee bee4324.550.473.655.45
Flexural strength4326.689.530.0026.40
Split tensile strength4321.422.020.005.16
Compressive strength43241.917.5021.2859.75

Statistical summary of input and output variables.

4.1 Feature correlation

Feature correlation of the data set was conducted to determine the highly correlated input parameter or variable with the target variable or strength. Figure 1 shows the feature correlation heat map of the dataset. The correlation map shows that cement, superplasticizer, age, flexural strength, and split tensile strength are highly correlated with compressive strength; age is a major input parameter affecting it, followed by cement and superplasticizer. The split tensile strength is highly correlated to age, flexural strength, and compressive strength, while flexural strength is in turn correlated to age, split tensile strength, and compressive strength.

FIGURE 1

4.2 Performance metrics

The dataset was further divided into training and testing data in the ML technique. To model the flexural, split tensile, and compressive strengths of SFRC, 432 samples were randomly divided into 345 samples (80%) for the training process and 87 samples (20%) for the testing process. The nine ML methods (KNN, RF, SVR, DT, GB, AB, XGBoost, CATBOOST, and LGBM) were organized to project the flexural, split tensile, and compressive strengths of SFRC. The dataset was thoroughly evaluated by the metrics coefficient of determination (R2) and mean squared error (MSE). Table 10 summarizes the values obtained. It can be observed from there that the ensemble-based algorithms had better prediction than the single learner models. The R2 values of KNN were found to be lowest for all three strengths, while SVR, RF, and DT were found to be close to predicting the split tensile strength. The booting ensemble methods—GB, CATBOOST, AB, LGBM, and XGBR—provided excellent R2 values for predicting flexural and split tensile strengths. However, the CAT BOOST algorithm was found to have overfitting values. The R2 values for LGBM and XGBR for the compressive strength indicated good model predictions, with the latter being superior in predicting all three strengths. The R2 value for the XGBR algorithm was found to be 0.9926, 0.9965, and 0.7837 for the flexural, split tensile, and compressive strengths, respectively. The results show that boosting-based algorithms better capture the nonlinear relationships that exist between the input parameters and the mechanical behavior of SFRC through highly accurate modeling of the flexural behavior dependent on the distribution of fiber.

TABLE 10

Sl NoAlgorithmSplit tensile Flexural Compressive Split tensile MSEFlexural MSECompressive MSE
1KNN0.0930.09170.09473.826685.486851.8273
2SVR0.96360.79850.65520.153818.962619.7362
3RF0.96360.79850.65520.153818.962619.7362
4DT0.96360.79850.65520.153818.962619.7362
5GB0.99410.98380.54210.02491.524726.2151
6CatBoost0.99640.99120.68130.01510.832818.2422
7AdaBoost0.99640.99180.67950.01510.769718.3506
8LGBM0.99650.99250.73580.01490.707913.8530
9XGBoost0.99650.99260.78370.01470.698012.3828

Performance comparison of machine learning models.

The MSE values also had a similar impact, which showed that the XGBR ensemble boosting algorithm provided an excellent model prediction compared to other ML methods for all three strengths. The second in line to it was the LightGBM method. Through these results, it is clear that the prediction performance of ensemble-based boosting algorithms is superior and, therefore, ideal for NDT-based prediction in SFRC and thus better suited for applications related to SHM.

There have been several studies involving ML in SFRC, but they concentrated on a single target property, primarily compressive strength. However, we have used the ML models to predict split tensile, flexural, and compressive strengths, thus evaluating the performance of SFRC more effectively. This study systematically compared nine different ML models using the same preprocessing and training conditions, thus ensuring a fair and unbiased comparison. It is also linked to real-world monitoring scenarios rather than only laboratory-based predictions and has also resulted in the creation of a novel dataset which can be used later for further predictions.

4.3 Training vs. testing performance

To rule out overfitting, a comparison of the training and testing performance was conducted for all three strengths separately; the results of the comparison for split tensile strength are tabulated in Table 11, while results for flexural and compressive strengths are shown in Tables 12 and 13, respectively.

TABLE 11

Algorithm (training) (test)MSE (training)MSE (test)
KNN0.14000.0933.193.8266
SVR0.97930.96360.13160.1538
RF0.98170.96360.12100.1538
DT0.98310.96360.11900.1538
GB0.99570.99410.01900.0249
CatBoost0.99700.99640.01710.0151
AdaBoost0.99670.99640.01800.0151
LGBM0.99690.99650.01590.0149
XGBoost0.99720.99650.01440.0147

Performance comparison of models for split tensile strength.

TABLE 12

Algorithm (training) (test)MSE (training)MSE (test)
KNN0.14350.091775.3485.4868
SVR0.82410.798516.8718.9626
RF0.82000.798516.6018.9626
DT0.83120.798516.8018.9626
GB0.99600.98381.231.5247
CatBoost0.99570.99120.960.8328
AdaBoost0.99600.99180.840.7697
LGBM0.99680.99250.830.7079
XGBoost0.99710.99260.760.6980

Performance comparison of models for flexural strength.

TABLE 13

Algorithm (training) (test)MSE (training)MSE (test)
KNN0.09650.094745.2851.8273
SVR0.72120.655218.5419.7362
RF0.72430.655218.3319.7362
DT0.72700.655218.4619.7362
GB0.61000.542124.6026.2151
CatBoost0.74410.681317.8718.2422
AdaBoost0.73700.679518.1418.3506
LGBM0.81000.735812.5313.8530
XGBoost0.83420.783711.2712.3828

Performance comparison of models for compressive strength.

From these tables, it is evident that the difference is minimal in the training and testing performance of most models. This indicates a good generalization capability. It can also be observed that ensemble-based models such as XGBoost, LightGBM, and CatBoost have better performance consistently across all strength parameters.

4.4 Model interpretability and engineering insights

The model’s interpretability was analyzed by performing feature extraction analysis on the developed ensemble ML models. Parameters such as the cement content, fiber content, and water–cement ratio strongly affect the prediction of mechanical properties. It was observed that fiber content positively contributed to split tensile and flexural strengths due to crack-bridging mechanisms. The lower water–cement ratio enhances strength. SHAP (Shapley Additive exPlanations) analysis provided a quantitative measure of the contribution of each variable to the model performance. The SHAP plot is shown in Figure 2. These results indicate that a critical role is played by the interaction between the material composition parameters in predicting the strength properties. This is due to the nonlinear behavior identified by the ML models. This shows that the ensemble models have the ability to model complex relationships in SFRC. It can be observed from the tabulated results that compressive strength prediction has a lower performance, with a R2 value of 0.7837, compared to tensile and flexural strengths. This can be attributed to the fact that microstructural and material heterogeneities largely affect the compressive strength. These may include conditions used for curing, distribution pattern of aggregates, and voids which may result in dataset inconsistency. Fiber reinforcement mechanisms are prominent in tensile and flexural strengths and hence were captured more effectively by the input parameters. In addition, the brittle and non-linear nature of compressive failure makes it more difficult for the ML models to generalize accurately.

FIGURE 2

4.5 Practical implementation in NDT-based structural health monitoring systems

The framework proposed in this study can be integrated into SHM systems by incorporating data acquired using NDT with predictive data analytics for the real time assessment of concrete strength. Figure 3 illutrates the machine learning based SHM framework for SFRC for practical implementation. NDT techniques that are commonly used, such as acoustic emission monitoring, ultrasonic pulse velocity, and rebound hammer testing, can be used to capture the response of the material without causing any destruction or damage to the structure. By using these measurements with the mix proportion design parameters as the inputs to the trained ML models, the tensile, flexural, and compressive strengths of concrete can be predicted. The methodology includes a sequential pipeline that consists of data acquisition and preprocessing, model inference, and decision-making support. The NDT data collected are first filtered to remove noise, and then the inputs are normalized. The trained ML models are then applied on these inputs either on-site or through a cloud-based platform which can provide the strength predictions as output. These predicted values can support engineers to detect damage, assess conditions, and plan maintenance. In real-world scenarios, environmental factors like moisture, temperature, and aging effects can influence NDT measurements. Moreover, since the trained model consists of controlled experimental data, it may not perform consistently in diverse field scenarios. The prediction accuracy is also affected by the uncertain knowledge of mix proportions. This requires a periodic calibration of the model using field data which can significantly improve performance. Thus, the proposed study provides an efficient and scalable approach for including data-driven models in SHM systems, thus ensuring a faster and non-invasive assessment of concrete.

FIGURE 3

5 Conclusion

This study demonstrates that the incorporation of steel fibers significantly enhances the mechanical performance of concrete, with optimal improvements observed up to a fiber content of 1.5%, beyond which a decline in strength was noted. Both tensile-related and compressive properties exhibited similar trends, indicating the existence of an optimum fiber dosage for effective stress transfer and crack control. Parametric analysis revealed that age, cement content, and superplasticizer dosage strongly influence compressive strength, whereas age emerged as the dominant factor governing flexural and split tensile strengths. Among the machine learning models evaluated, model performance was primarily assessed using the coefficient of determination (R2) alongside error metrics to ensure robustness. Although the CATBOOST algorithm yielded high R2 values for compressive strength prediction, its susceptibility to overfitting limited its reliability. Overall, the XGBoost regressor demonstrated superior predictive capability and generalization, achieving R2 values of 0.9926, 0.9965, and 0.7837 for flexural, split tensile, and compressive strengths, respectively, and is, therefore, identified as the most effective model for predicting the mechanical behavior of steel fiber reinforced concrete.

6 Future scope

Future research may extend the evaluation of mechanical properties beyond the 28-day curing period to capture the long-term performance and durability of reinforced steel fiber and high-performance concretes. Additional mechanical parameters, including modulus of rupture and Poisson’s ratio, should be investigated to provide a more comprehensive understanding of structural behavior. The study may also be expanded to include other performance-related tests relevant to high-performance concrete applications. From a modeling perspective, the use of additional performance metrics is recommended to further assess and compare predictive robustness. Improvements to the CATBOOST model, particularly strategies to mitigate overfitting, warrant further investigation. Moreover, the application of deep-learning-based approaches presents a promising direction for enhancing prediction accuracy and capturing complex nonlinear relationships in concrete behavior.

Statements

Data availability statement

The raw data supporting the conclusions of this article will be made available by the authors without undue reservation.

Author contributions

PS: Writing – original draft, Formal analysis, Funding acquisition, Writing – review and editing, Project administration, Visualization, Software, Methodology, Data curation, Resources, Conceptualization, Investigation, Validation. RR: Supervision, Conceptualization, Writing – review and editing, Writing – original draft, Project administration, Validation. RD: Methodology, Writing – original draft, Supervision, Writing – review and editing, Conceptualization. PN: Writing – original draft, Conceptualization, Validation, Writing – review and editing, Supervision, Methodology. SD: Software, Visualization, Validation, Writing – original draft, Writing – review and editing.

Funding

The author(s) declared that financial support was not received for this work and/or its publication.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that generative AI was not used in the creation of this manuscript.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Abbreviations

KNN, K-nearest neighbors; DT, decision tree; SVR, support vector regression; RF, random forest; GB, gradient boosting; AB, AdaBoost; XGBoost, extreme gradient boosting; LGBM, light gradient boosting machine; CatBoost, categorical boosting; SFRC, steel fiber reinforced concrete; MSE, mean squared error; ML, machine learning; ANN, artificial neural networks; SVM, support vector machines; SHM, structural health monitoring; NDT, non-destructive testing; UPV, ultrasonic pulse velocity; SHAP, Shapley additive explanations.

References

  • 1

    Al-DouriA. R.Al-JanabiM. A.AhmedS. S. (2024). Comparative evaluation of k-nearest neighbor and ensemble learning models for engineering datasets. Expert Syst. Appl.235, 121028.

  • 2

    Al-ShamasnehA. R.MahmoodzadehA.KarimF. K.SaidaniT.AlghamdiA.AlnahasJ.et al (2025). Application of machine learning techniques to predict the compressive strength of steel fiber reinforced concrete. Sci. Rep.15, 30674. 10.1038/s41598-025-16516-1

  • 3

    AswalV.SinghB.MaheshwariR. (2024). Machine learning-based model for prediction of concrete strength. Multiscale Multidiscip. Model. Exp. Des.8, 48. 10.1007/s41939-024-00609-x

  • 4

    AwolusiT. F.OkeO. L.AkinkurolereO. O.SojobiA. O.AlukoO. G. (2019). Performance comparison of neural network training algorithms in the modeling properties of steel fiber reinforced concrete. Heliyon5, e01115. 10.1016/j.heliyon.2018.e01115

  • 5

    BanuT. U.RajamaneN. P.AwoyeraP. O.GobinathR. (2020). Strength characterisation of self-cured concrete using ai tools. Mater. Today Proc.39, 839848. 10.1016/j.matpr.2020.10.101

  • 6

    Bureau of Indian Standards (1959). IS 516:1959 – method of tests for strength of concrete. New Delhi, India: Bureau of Indian Standards.

  • 7

    Bureau of Indian Standards (1988). IS 4031:1988 – methods of physical tests for hydraulic cement. New Delhi, India: Bureau of Indian Standards.

  • 8

    Bureau of Indian Standards (1999). IS 5816:1999 – method of test for splitting tensile strength of concrete. New Delhi, India: Bureau of Indian Standards.

  • 9

    Bureau of Indian Standards (2000). IS 456:2000 – code of practice for plain and reinforced concrete. New Delhi, India: Bureau of Indian Standards.

  • 10

    Bureau of Indian Standards (2009). IS 10262:2009 – Concrete mix proportioning – guidelines. New Delhi, India: Bureau of Indian Standards.

  • 11

    CaldaroneM. A. (2019). High-strength concrete: a practical guide. 1 edn.Boca Raton, FL, USA: CRC Press.

  • 12

    ChakmaJ.ZhouZ.ChakmaB. (2025). Mechanical strength prediction of steel-polypropylene fiber-based high-performance concrete using hybrid machine learning algorithms. arXiv. 10.48550/arXiv.2512.21638

  • 13

    ChoiD.HongK.OchirbudM.MeiramovD.SukontaskuulP. (2023). Mechanical properties of ultra-high-performance concrete and ultra-high-performance fiber-reinforced concrete with recycled sand. Int. J. Concr. Struct. Mater.17, 67. 10.1186/s40069-023-00631-2

  • 14

    DaoD. V.AdeliH.LyH. B.LeL. M.LeV. M.LeT. T.et al (2020). A sensitivity and robustness analysis of gpr and ann for high-performance concrete compressive strength prediction using a monte carlo simulation. Sustainability12, 830. 10.3390/su12030830

  • 15

    El-SayedM.IbrahimA.HassanH. A. (2024). Support vector regression-based modelling for nonlinear prediction in engineering applications. Eng. Appl. Artif. Intell.130, 106945.

  • 16

    ElhishiS.ElashryA. M.El-MetwallyS. (2023). Unboxing machine learning models for concrete strength prediction using xai. Sci. Rep.13, 19892. 10.1038/s41598-023-47169-7

  • 17

    FuH.ZhouX.XuP.SunD. (2025). Prediction of compressive strength of concrete using explainable machine learning models. Materials18, 5009. 10.3390/ma18215009

  • 18

    GouJ.ZamanA.FarooqF. (2025). Machine learning-based prediction of compressive and split tensile strength of steel fiber-reinforced recycled aggregate concrete. Eng. Appl. Artif. Intell. 161. 10.1016/j.engappai.2025.112190

  • 19

    IleriK. (2025a). Comparative analysis of catboost, lightgbm, xgboost, random forest, and decision tree methods optimized with particle swarm optimization. Int. J. Mach. Learn. Cybern.16, 69376956.

  • 20

    IleriK. (2025b). Catboost-based optimized learning framework for high-accuracy prediction. Int. J. Mach. Learn. Cybern.16, 69376956. 10.1007/s13042-025-02654-5

  • 21

    KangM. C.YooD. Y.GuptaR. (2021). Machine learning-based prediction for compressive and flexural strengths of steel fiber-reinforced concrete. Constr. Build. Mater.266, 121117. 10.1016/j.conbuildmat.2020.121117

  • 22

    KaviyaK.PremalathaJ. (2019). Prediction of compressive strength of high-performance concrete using ann. Int. Res. J. Eng. Technol.6, 13781387.

  • 23

    KhanM. A.GandomiA. H.MirjaliliS. (2024). Extreme gradient boosting for high-dimensional regression problems: a comparative study. Knowledge-Based Syst.292, 111660.

  • 24

    KumarS.MishraR. K.SamuiP. (2024). Random forest-based prediction and uncertainty analysis for civil engineering datasets. Automation Constr.162, 105305.

  • 25

    LiuY. (2022). High-performance concrete strength prediction based on machine learning. Comput. Intell. Neurosci.2022, 5802217. 10.1155/2022/5802217

  • 26

    LiuK.ZhangL.WangW.ZhangG.XuL.FanD.et al (2023). Development of compressive strength prediction platform for concrete materials based on machine learning techniques. J. Build. Eng.80, 107977. 10.1016/j.jobe.2023.107977

  • 27

    LiuS.CaoS.HaoY.ChenP.MaG. (2024). Prediction models of compressive mechanical properties of steel fiber-reinforced cementitious composites. J. Build. Eng.84, 108629. 10.1016/j.jobe.2024.108629

  • 28

    MahmoudM. R.El-SheikhA.Abdel-WahabT. (2024). Performance evaluation of gradient boosting models for regression problems in engineering. Appl. Soft Comput.150, 111038.

  • 29

    MarvilaM. T.de AzevedoA. R. G.de MatosP. R.MonteiroS. N.VieiraC. M. F. (2021). Materials for production of high and ultra-high-performance concrete: review and perspective of possible novel materials. Materials14, 4304. 10.3390/ma14154304

  • 30

    MuhammadU. J.AminuI. I.MahmoudI. A.AliyuU. U.UsmanA. G.JibrilM. M.et al (2024). An improved prediction of high-performance concrete compressive strength using ensemble models and neural networks. AI Civ. Eng.3, 21. 10.1007/s43503-024-00040-8

  • 31

    MurtiningsihD. A.SariB. W.FajriI. N. (2025). Comparison of lightgbm, xgboost, and catboost with balancing and hyperparameter tuning. J. Appl. Inf. Comput.9, 27532763. 10.30871/jaic.v9i5.10400

  • 32

    NevilleA.AïtcinP. C. (1998). High-performance Concrete–an overview. Mater. Struct.31, 111117. 10.1007/bf02486473

  • 33

    NguyenT. T.DuyH. P.ThanhT. P.VuH. H. (2020). Compressive strength evaluation of fiber-reinforced high-strength self-compacting concrete with artificial intelligence. Adv. Civ. Eng.2020, 3012139. 10.1155/2020/3012139

  • 34

    PremkumarG.Senthil SelvanS. (2025). Prediction of mechanical properties of steel fiber-reinforced concrete under elevated temperature using artificial neural networks. Front. Built Environ. 11. 10.3389/fbuil.2025.1610115

  • 35

    ShrivastavaS.ShrivastavaT. (2024). Prediction of concrete’s compressive strength using machine learning algorithms. Mater. Today Proc.103, 183189. 10.1016/j.matpr.2023.08.252

  • 36

    SobuzM. R.AdittoF. S.DattaS. D.KabboM. K. I.JabinJ. A.HasanN. M. S.et al (2024). High-strength self-compacting concrete production incorporating supplementary cementitious materials: experimental evaluations and machine learning modelling. Int. J. Concr. Struct. Mater.18, 67. 10.1186/s40069-024-00707-7

  • 37

    SuN.GuoS.ShiC.ZhuD. (2024). Predictions of mechanical properties of fiber reinforced concrete using ensemble learning. J. Build. Eng. 98. 10.1016/j.jobe.2024.110990

  • 38

    VihasC.Rama RaoP.Santosh KumarD. L. M.ReddyV. V. S.AbhilashN. (2025). Machine learning-based prediction of concrete compressive strength incorporating ggbs. Procedia Struct. Integr.70, 461468. 10.1016/j.prostr.2025.07.078

  • 39

    WangH.LinJ.GuoS. (2025). Study on the compressive strength predicting of steel fiber reinforced concrete based on an interpretable deep learning method. Appl. Sci. 15. 10.3390/app15126848

  • 40

    WijesundaraS.WijesundaraK.BandaraS. (2025). Machine learning approach for predicting the compressive strength of ultra-high-performance fiber reinforced concrete. Structures75, 108704. 10.1016/j.istruc.2025.108704

  • 41

    YanK.XuH.ShenG.LiuP. (2013). Prediction of splitting tensile strength from cylinder compressive strength of concrete by support vector machine. Adv. Mater. Sci. Eng.2013, 597257. 10.1155/2013/597257

  • 42

    YangY.LiuG.ZhangH.ZhangY.YangX. (2024). Predicting the compressive strength of environmentally friendly concrete using multiple machine learning algorithms. Buildings14, 190. 10.3390/buildings14010190

  • 43

    YazıcıS.InanG.TabakV. (2007). Effect of aspect ratio and volume fraction of steel fiber on the mechanical properties of sfrc. Constr. Build. Mater.21, 12501253. 10.1016/j.conbuildmat.2006.05.025

  • 44

    YennimarY.LeonardiW.WeideH.CantonaD.HutagalungG. M. (2024). Comparison of data mining algorithms based on adaptive boosting in predicting diabetes mellitus. J. Tek. Inform. C.I.T Medicom16, 112. 10.35335/cit.Vol16.2024.730.pp1-12

  • 45

    ZhangY.YeY.WangJ.TangB.FuF. (2025). Strength prediction of self-compacting concrete using improved rvm machine learning method. Int. J. Concr. Struct. Mater.19, 101. 10.1186/s40069-025-00835-8

  • 46

    ZhengD.WuR.SufianM.KahlaN. B.AtigM.DeifallaA. F.et al (2022). Flexural strength prediction of steel fiber-reinforced concrete using artificial intelligence. Materials15, 5194. 10.3390/ma15155194

Summary

Keywords

ensemble machine learning models, non-destructive testing, single target machine learning models, steel fiber reinforced concrete, strength prediction, structural health monitoring

Citation

Shijin PJ, Raghunandan Kumar R, Davis R, Nagarajan P and Das S (2026) Comparative evaluation of supervised machine learning models in the non-destructive strength prediction of steel fiber reinforced concrete. Front. Built Environ. 12:1811594. doi: 10.3389/fbuil.2026.1811594

Received

15 February 2026

Revised

29 March 2026

Accepted

01 April 2026

Published

12 May 2026

Volume

12 - 2026

Edited by

Emilio Bastidas-Arteaga, Université de la Rochelle, France

Reviewed by

Nishant Raj Kapoor, Academy of Scientific and Innovative Research (AcSIR), India

Rohit Maheshwari, DIT University, India

Updates

Copyright

*Correspondence: P. J. Shijin,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics