Abstract
Introduction:
Predicting change order magnitude in fragmented public construction systems remains methodologically unreliable, particularly under extreme data scarcity and correlated project attributes. This study develops a governance-calibrated predictive architecture for estimating the Change Order Ratio (COR) using objective contractual and design records from 21 public projects in Jordan, and is therefore framed as preliminary, proof-of-concept evidence from a small, single-context sample rather than externally generalizable inference.
Methods:
To prevent overfitting and validation bias, a leakage-free nested Leave-One-Out Cross-Validation (LOOCV) framework integrates regularized regression and nonlinear Support Vector Regression (SVR).
Results:
The unregularized baseline achieves strong in-sample fit (R2 of 0.84) but deteriorates under cross-validation (R2 of 0.54), suggesting instability in small datasets. Regularization enhances robustness (LASSO LOOCV R2 of 0.70), while nonlinear SVR delivers the strongest generalization (LOOCV R2 of 0.76; RMSE of 1.25), suggesting curvature and interaction structures beyond linear shrinkage in this sample. Sensitivity analysis identifies nonlinear amplification of design completeness and positions Requests for Information as leading instability indicators.
Discussion:
The findings suggest that change-order volatility is more closely associated with tender-stage information and documentation conditions than with project scale or contractual strictness in the studied context. The study contributes a defensible small-sample modeling framework and highlights upstream documentation maturity as the primary leverage point for enhancing predictability in data-scarce public-sector project delivery environments.
1 Introduction
Public sector investments in construction activities in developing countries, such as roads, education, and healthcare, remain a key focus for economic and social development. These investments are often subject to contractual changes in the post-award stage, including changes in scope, design, schedule, and quality improvements in construction activities (; ) Empirical studies in Jordan indicate a strong association between change orders and cost overruns, project delays, and design changes in public sector construction contracts (; ; ). Similar patterns have been reported in road construction (), risk forecasting () and regional cost estimation (; ), which overall suggests that change-related uncertainty remains a primary source of instability for the performance of construction projects.
The persistence of change orders is well documented, and the literature has extensively analyzed the causes, consequences, and control mechanisms of change orders. However, the predictability of change orders is considered a different issue. Change orders are not the product of any one cause. Instead, they are the complex result of technical, managerial, and institutional factors such as insufficient design documentation, delayed revisions, unforeseen geotechnical conditions, owner-initiated changes, regulatory changes, and unclear tender documents (; ). In the context of public projects in Jordan, these drivers are intensified by the inflexibility of the process, fragmented coordination, and multi-level procurement, all of which increase the probability of post-award modifications (; ). Similarly, in other countries, research has identified the relationship between poor design integration, poor project leadership, and poor control systems and the occurrence of change orders, including their frequency and magnitude (; ). Therefore, the ability to foresee change is now critical to the stability of project delivery, as changes impact workflow, including cost, time, and quality outcomes.
Nevertheless, the availability of such sophisticated management systems does not always mean that predictive capabilities can be achieved. In order to achieve quantitative prediction of COR, historical data sets must be consistent, the magnitude of change have to be accurately measured, and feature definitions must be similar across projects. In many public systems found in developing countries, project data can be scattered across organizations, unevenly computerized, and often exists in non-centralized paper-based systems. This limits the ability to develop consistent, usable data sets that can be used for econometric modeling. In Jordanian public systems, most studies have focused on identifying causes and impacts of variation orders, while the prediction of normalized metrics such as the COR from objective contractual records has been limited (; ).
In order to address this gap, the current study aims to develop a predictive framework specifically suited to extreme data-scarcity environments. By drawing upon objective contractual and design documentation from public construction projects in Jordan, the current study highlights a leakage-free modeling pipeline, wherein regularized regression methods are combined with nonlinear SVR and are cross-validated through a nested hyperparameter optimization and LOOCV approach. Not only does the current study aim to propose a methodologically defensible approach to COR prediction, but it also outlines a generalized architectural structure for enhancing governance-oriented decision-making in change-order risk environments in public construction delivery systems. Beyond the methodological contribution, the motivation for this work is practical. In data-scarce public construction systems, change orders are typically managed reactively and based on qualitative judgment, with little quantitative support at the tender stage. By linking objective tender documentation to expected COR under a leakage-free small-sample pipeline, the proposed framework can help public owners identify high-risk tenders earlier and prioritize design reviews, investigations, and contingency planning before contract award.
2 Literature review
Change orders in the context of construction projects have been recognized as a complicated phenomenon resulting from a multitude of interrelated technical, contractual, organizational, and environmental influences. From all the literature available, incomplete integration of designs, lack of effective project control systems, poor leadership coordination, and regulatory uncertainty have been identified to be the principal influences in regard to change orders (; ). Within the Jordanian and other public sectors, this issue is further exacerbated due to the lack of flexibility in procedures and fragmented procurement processes, that inherently increases the risks of post-award changes to contracts (; ). Although these studies provide significant diagnostic insights, they are descriptive in nature, seeking to identify the causative factors rather than developing predictive models to predict the extent of change orders in such environments.
The existing approaches of predicting the occurrence of change orders were mostly based on expert judgment, qualitative risk management, and classical statistical techniques (; ). Expert-based methodologies are known for their subjective character and low reproducibility (), While linear statistical modeling postulates that the relationship between different variables can be modeled using linear equations, which, in most cases, does not capture the nonlinear nature of the building systems. For example, the combined effect of design completeness and geotechnical risk on cost growth or change orders is often multiplicative and threshold-based—projects with both low design maturity and unfavorable ground conditions can experience a disproportionate escalation in variations compared with what a simple additive linear model would predict. In the case of data-scarce environments, the quality of the historical records, the inconsistency of the information, and the lack of continuity in the archival process can also question the reliability of the statistical results (). This challenge has frequently appeared in construction research from various regions of the Middle East where project data can be dispersed across agencies, inconsistently digitized, or not systematically archived. Therefore, robust statistical modeling and acceptable accuracy prediction is very difficult. In these cases, conventional methods identify only symptoms of causes that have already materialized after change orders have occurred.
Machine learning (ML) provides a wide range of modeling approaches to address nonlinear and high-dimensional relationships effectively (; ). ML models, including SVR (), Artificial Neural Networks (; ), Random Forests (), and many other techniques have suggest strong predictive performance in cost estimation (), schedule prediction (), risk/safety evaluation (), quality assessment (), dispute prediction (), and risk classification using explainable models (). Nonetheless, the majority of the applications of ML methods assume the availability of large enough structured sets of data. However, in a fragmented public sector with limited instances and heterogeneous sets of projects, the dangers of overfitting, parameter estimation instability, and overreporting of accuracy tend to be significantly amplified (; ). Overcoming these challenges requires the adoption of specific strategies for small sets of data. For example, the application of LASSO and Ridge regressions help to overcome the challenges of instability by reducing the complexity of the models (). In addition, the adoption of strict validation strategies, including LOOCV and the optimization of the hyperparameters, help to overcome the challenges of leakage and defensible evaluation of the process of generalization (; ). These strategies are increasingly being advocated for the application of regional construction analytics under limited sets of data (; ).
In the context of this research, these techniques are not used only as technical tools to fit a model on a small dataset, but as a way to obtain more stable and defensible signals about which tender-stage factors consistently influence COR. By reducing overfitting and variance inflation, the regularized models and strict validation allow us to distinguish robust governance-related effects (for example, documentation completeness and RFIs) from patterns that are likely to be artefacts of a small and heterogeneous sample.
Apart from the algorithm, there is a data origin gap in terms of structure. Most of the research in predictive analysis is developed their prediction based on data or datasets obtained from surveys or aggregate sources rather than actual execution-stage contractual records (; ). Surveys are useful in perceptual analysis, although they are open to data response bias and may lack sufficient technical detail (). Contractual records, on the other hand, are a direct expression of the administrative and technical context of change formation (; ). In spite of this distinction, the predictive estimation of normalized metrics like the COR based on objective execution records faces limited capabilities in the context of scarce data environments such as the Jordanian public sector (; ). In addition, many ML-based research efforts fail to address the estimation of change orders as a problem of small-sample inference with strict control over the separation of hyperparameter tuning and performance evaluation (; ). In the context of limited and heterogeneous datasets, this might lead to overestimated accuracy at the cost of generalization capabilities (; ). The lack of a rigorously regularized estimation framework for the magnitude of change orders based on objective contract records represents a clear research gap.
The current research seeks to address this issue by proposing a prediction model specifically designed for data-scarce public construction settings. A COR estimation model is developed based on objective contract and design documents from public projects in Jordan. A nonlinear architecture of SVR with embedded regularized regression and nested hyperparameter optimization/LOOCV validation is utilized to stabilize the estimation process and address the issue of overfitting under data-scarce conditions. Objective data from project records instead of perception-based surveys provide better connection between the input of the prediction model and the processes of change generation, and give an easily transferable methodology for public construction systems based on a combined set of data systems.
In contrast to most previous work, which is primarily descriptive or survey-based, the present study focuses on tender-stage, quantitative prediction of COR using objective contract and design records under extreme data scarcity. The contribution is therefore twofold: first, a leakage-free small-sample pipeline that can be reused in similar public data environments, and second, quantitative evidence that information-governance variables (design completeness and RFIs) exert stronger and more nonlinear influence on COR than project size alone.
3 Methodology
3.1 Data and variable construction
The data set contains 21 public sector construction projects. These projects are obtained through official records of the Jordanian government. The data variables include objective financial, technical, design, and geotechnical characteristics. Six predictors were included: Tender Value per square meter (TV), Design Completeness at tender stage (DC), Soil/Geotechnical Risk Level (GRL), Project Duration (DUR), Complexity Level (CL), and Requests for Information (RFIs). The variables are associated with project scale, design completeness, soil/geotechnical risk, project duration, complexity level, and requests for information.
Design Completeness (DC) was quantified as the proportion of design documentation completed and approved at tender stage, based on a standardized checklist of architectural, structural, mechanical, electrical, and geotechnical packages. Complexity Level (CL) was assigned on a three-point ordinal scale using pre-defined criteria that considered the number of functional systems, architectural and structural irregularity, and the presence of special technologies or non-standard elements. Soil/Geotechnical Risk Level (GRL) was coded on a three-point ordinal scale (1–3) reflecting increasing uncertainty and difficulty, based on the available geotechnical investigations and foundation design notes.
GRL and CL were independently coded by two members of the research team with civil engineering backgrounds. The coders were not provided with the final COR values during the coding process, although they could see the general project descriptions and contract files. Minor discrepancies between the two coders were resolved by discussion, and a high level of agreement was reached across projects (no formal kappa statistic was calculated, which we acknowledge as a limitation). RFIs were used as raw counts of documented design issues per project; they were not normalized by project size or documentation volume, and this simplification is also noted as a limitation for future work.
GRL was encoded as an ordered three-level classification reflecting the increasing geotechnical uncertainty: level 1 (stable conditions), level 2 (moderate challenges), and level 3 (high-risk conditions such as expansive soils, karst, or unknown utilities). Similarly, CL was encoded as an ordered classification structure. Continuous variables were standardized in each training fold to account for scale invariance, and the ordinal encodings were preserved. The response variable (COR) is defined as the percentage of the total approved change order amount with respect to the value of the awarded contract. Descriptive statistics for all variables are presented in Table 1.
TABLE 1
| Variable name | Symbol | Unit | Description | Min | Max | Mean | Std. Dev |
|---|---|---|---|---|---|---|---|
| Tender Value at Award per Square meter | TV | JD/m2 | Awarded unit cost of the project per square meter | 151.37 | 1149.63 | 583.32 | 236.04 |
| Design completeness percentage at tender | DC | (%) | Percentage of design completed at tender stage | 0.80 | 0.99 | 0.9133 | 0.0639 |
| Number of RFIs/Design errors | RFIs | Count | Total number of design issues or RFIs submitted | 2 | 25 | 6.95 | 6.84 |
| Soil/Geotechnical risk level | GRL | Level (1–3) | Coded geotechnical risk classification | 1 | 3 | 2.524 | 0.680 |
| Project duration | DUR | Months | Planned project duration | 12 | 51 | 36.10 | 12.37 |
| Complexity level | CL | Level (1–3) | Coded project complexity classification | 1 | 3 | 1.714 | 0.784 |
| Change orders ratio | COR | % | Actual total change orders divided by tender award value | 0.057 | 9.996 | 2.585 | 3.085 |
Summary of the model variables.
3.2 Model architecture and comparative
The selection of models was constrained by two conditions including extreme sample size constraint, and the need for stable and interpretable predictions. In addition, the predictors are correlated and the relationship between tender-stage variables and COR is plausibly nonlinear (e.g., threshold and interaction effects), which can make unregularized linear models unstable and poorly generalizable in such a small sample. For this reason, we combine regularized linear models with a nonlinear SVR specification to balance variance control and flexible response surfaces. A Multiple Linear Regression (MLR) model was defined as the baseline model (un-regularized model). Three supervised models are selected in order to predict COR including (1) LASSO Regression (L1-regularized linear model), (2) Ridge regression (L2-regularized linear model), and SVR with Radial Basis Function (RBF) kernel. Both LASSO and Ridge models aim to control the variance of the estimator by penalization, while the SVR model is used to achieve a nonlinear approximation of the response surface. All models are trained on the same set of predictors and evaluated according to a unified validation criterion. The predictive performance was assessed by the use of the Mean Absolute Error (MAE), the Root Mean Squared Error (RMSE), and the coefficient of determination (R2). The analysis was carried out in Python 3.11. The libraries used included scikit-learn 1.4 and numpy 1.26.
3.3 Regularized linear models
LASSO was chosen based on its appropriateness for small sample estimation and its capacity for simultaneous shrinkage and selection. The estimator is given by:where is the regularization parameter. The use of the -penalty term induces sparsity, as small values of the regression coefficients are driven to zero, thus improving estimation stability with small sample sizes.
On the other hand, Ridge regression tackles the issue of multicollinearity with the use of the -regularization term:
The penalty term is included to improve the conditioning of the Gram matrix and to reduce variance, especially with the presence of correlation between cost, duration, and design variables.
3.4 Support vector regression
To address the possible nonlinear relationship between the predictor variables and COR, ε-insensitive SVR with an RBF kernel was applied (; ). In the primal form, the optimization problem can be stated as follows:subject to ε-insensitive constraints. In the dual form, the optimization problem can be stated as follows:with RBF kernel:
Hyperparameters jointly regulate model complexity, margin tolerance, and surface smoothness.
3.5 Training and validation strategy
In consideration of the critical sample constraint, the model’s performance was evaluated using a completely nested and leakage-free framework of LOOCV (; ). For each observation , the model was trained on the remaining projects, while the held-out project was reserved for testing. Within each LOOCV, hyperparameters were tuned using an inner 3-fold cross-validation procedure.
To avoid information leakage, all numeric features were standardized using parameters estimated on the training set:where and represent the mean and standard deviation estimated on the respective data set of the training fold. The optimized model is further refitted on the scaled data and utilized to compute the prediction for the held-out observation (; ).
The mean squared error of LOOCV is given by:
The nested procedure has the advantage of separation between hyperparameter tuning and testing so that there is no optimistic bias and provides an unbiased measure of the model’s generalization ability. The final performance metrics (MAE, RMSE, and ) were estimated based on the aggregated LOOCV predictions. For all models, hyperparameters were tuned using an exhaustive grid search within the inner 3-fold cross-validation loop, with mean squared error (MSE) as the optimization criterion. For LASSO and Ridge regression, we used logarithmic grids of regularization strengths covering small to relatively large penalties (approximately from 10–4 up to 102). For the SVR with RBF kernel, we used a grid covering a wide range of model complexities, with C values from low to high (order of 10° to 103), ε values from coarse to very tight error tolerances (around 10–1 to 10–3), and γ values from very smooth to more flexible kernels (around 10–2 to 10°). The same grids were used in every outer LOOCV fold. When two or more hyperparameter settings produced essentially the same inner-CV MSE, we selected the simpler configuration (smaller penalty for LASSO/Ridge; smaller C and γ for SVR). Sensitivity curves for the SVR model were generated using a one-variable-at-a-time perturbation approach, where each predictor was varied over its observed range while holding the remaining predictors fixed, and the corresponding predicted COR values were plotted. This highly controlled validation pipeline guarantees a fair comparison of linear and nonlinear approaches while explicitly managing model complexity, parameter volatility, and validation bias. In a situation of data scarcity of this nature, such a controlled environment is necessary to ensure that the chosen SVR is a reflection of the genuine capabilities of the algorithm rather than validation artefacts.
4 Empirical results and model evaluation
4.1 Baseline linear model and overfitting diagnosis
The baseline MLR model was developed as an unregularized benchmark, connecting tender stage variables to COR as a dependent variable. As presented in Table 2, the model exhibits strong in-sample fit (R2 of 0.84; F of 12.65, p < 0.001) and low estimation errors (RMSE of 0.998; MAE of 0.752). The high degree of alignment between predicted and actual values in Figure 1a suggests that the model is successful in explaining variance within the estimation sample. Nonetheless, this apparent adequacy does not extend to the predictive evaluation. As seen in Table 2, when the nested LOOCV evaluation framework is used, the performance significantly deteriorates. In this case, the R2 value drops to 0.54, whereas the prediction errors increase (RMSE of 1.72, MAE of 1.24). This instability when extrapolating to unseen projects is also reflected by the greater spread of the points around the identity line in Figure 1b. The difference between the strong in-sample fit in Figure 1a and the weakened LOOCV alignment in Figure 1b provides evidence of overfitting. In the case of a very small sample size and correlated predictors, the unregularized model may fit sample-specific noise in addition to the underlying structural signal, thereby inflating in-sample explanatory power and weakening generalization to unseen samples. The magnitude of the performance gap indicates variance-driven instability and limited transferability across heterogeneous project cases.
TABLE 2
| Model performance metric | Value |
|---|---|
| In-sample R2 | 0.84 |
| F-statistic | 12.651 |
| F-statistic p-value | 0.000060 |
| Residual Standard error (RSE) | 1.2225 |
| Training RMSE | 0.9981 |
| Training MAE | 0.7524 |
| LOOCV R2 | 0.54 |
| LOOCV RMSE | 1.7175 |
| LOOCV MAE | 1.2412 |
Summary of baseline linear regression model performance.
FIGURE 1
This highlights a methodological baseline against which all subsequent modeling strategies must be measured. These results show that statistical significance and a large in-sample R2 are not sufficient as a proxy for predictive reliability under conditions of severe data scarcity. The baseline model serves as a diagnostic reference to illustrate the necessity of regularization and controlled model complexity for defensible levels of generalization performance.
4.2 Regularized linear models (stability and sparsity)
Regularization was developed to manage the variance of the estimators and ensure the stability of inference, especially in small sample sizes with issues of multicollinearity. The variables of construction tenders, including cost intensity, duration, maturity, and coordination, tend to have structural correlation patterns, which increase the variance of MLR estimation. The penalized regression directly manages the instability of estimation by shrinking the magnitude of the estimators.
As such, LASSO, optimized through a nested LOOCV at λ of 0.3044, is found to produce a sparse model wherein only the DC predictor is found to retain a non-zero coefficient weight. The other predictors are all eliminated through the penalization process, suggesting that DC is the only predictor that is able to identify linear relationships stable enough to pass the test of out-of-sample evaluation. As presented in Table 3, LASSO is able to maintain strong in-sample fit (R2 of 0.80) as well as improved generalization performance (LOOCV R2 of 0.70; RMSE of 1.40; MAE of 1.02). The well-defined minimum of the LOOCV MSE curve presented in Figure 2 suggests that the optimal penalty is neither arbitrary nor a result of bias-variance tradeoffs. The predictive ability of the sparse specification is shown in Figure 3 below. Figure 3a shows that the training fit continues to closely follow the identification line even after the model has been simplified to a single variable, which confirms our observation that model simplicity does not harm in-sample fit significantly. More interestingly, Figure 3b shows a tighter clustering around the identification line for the LOOCV estimate relative to the unregularized estimate. The induced sparsity has a structural interpretation. The exclusion of TV, RFIs, GRL, DUR, and CL suggests that their relative explanatory significance is not robust when controlling for variance inflation. On the other hand, DC remains robust to penalization, supporting its dominant linear role in explaining COR under constrained data conditions.
TABLE 3
| Performance metric | Value |
|---|---|
| In-sample R2 | 0.80 |
| Training RMSE | 1.138,028 |
| Training MAE | 0.860540 |
| LOOCV R2 | 0.70 |
| LOOCV RMSE | 1.395431 |
| LOOCV MAE | 1.018542 |
Summary of LASSO regression model performance.
FIGURE 2
FIGURE 3
On the other hand, the Ridge Regression is optimized at λ of 4.198, adopts a contrasting strategy of stabilization. The error curve of LOOCV presented in Figure 4 shows a flat curve for small values of λ, a minimum value of LOOCV error at λ of 4.198, and a sharp increase for larger values. This is indicative of bias dominance. The value of the penalty is thus a data-supported bias-variance tradeoff. Unlike the LASSO, Ridge Regression does not adopt sparsity as a strategy for dimensionality reduction. Instead, all coefficients are shrunk to achieve a distribution of the explanatory power of the independent variables among the correlated variables. As presented in Table 4, the performance of the Ridge Regression method is enhanced as compared to the baseline model without any regularization (LOOCV R2 of 0.60; RMSE of 1.59; MAE of 1.07), although it is less stable compared to the LASSO Regression. The prediction pattern of the Ridge Regression presented in Figure 5b shows a reduced level of variance as compared to the baseline model, although a large amount of dispersion is observed.
FIGURE 4
TABLE 4
| Performance metric | Value |
|---|---|
| In-sample R2 | 0.82 |
| Training RMSE | 1.078540 |
| Training MAE | 0.745624 |
| LOOCV R2 | 0.60 |
| LOOCV RMSE | 1.593901 |
| LOOCV MAE | 1.068692 |
Summary of Ridge regression model performance.
FIGURE 5
The difference between L1 and L2 penalty terms offers an interesting structural observation. Sparsity might be more protective against variance-related instabilities than smooth shrinkage alone in extreme data-scarce scenarios. Design maturity stands out as the sole linear effect that generalizes across projects, while secondary effects show weaker or interaction-related influences that are not yet fully stabilized within a linear additive model structure.
4.3 Nonlinear learning (SVR performance under nested validation)
In order to determine if the formation of COR is mediated by non-additive linear structures, SVR model with an RBF kernel was fitted under the same nested validation framework. The aim is not only to improve predictive accuracy, but also to test whether allowing controlled nonlinearity yields more stable out-of-sample behavior than purely linear shrinkage models in this data-scarce setting. The optimal parameter configuration was determined by the values of C of 1000, ε of 0.0005, and γ of 0.05, which were exclusively determined using a LOOCV-based search. The model suggest a nearly flawless training performance, as indicated by the high R2 value of 0.99 (Table 5), which is consistent with the close match in Figure 6a. However, the predictive sufficiency is measured by cross-validated behaviors. Under the nested LOOCV, SVR has the best generalization performance among all models, as shown in Table 5, with LOOCV R2 of 0.76, RMSE of 1.25, and MAE of 1.00, and Figure 6b shows a tight clustering around the identity line compared with other linear models. The surface for the hyperparameters in Figure 7 shows a well-defined area of low error rates, which are associated with large values of C and moderate values of γ. Large values of error occur when the smoothness constraint is too strong (too small a value of γ) or too rigid (too large a value of γ). Too small a value of C makes the margin constraint too weak. The selected values fall within a stable bias/variance regime indicated by the minimum of the LOOCV curve.
TABLE 5
| Performance metric | Value |
|---|---|
| In-sample R2 | 0.99 |
| Training RMSE | 0.000473 |
| Training MAE | 0.000406 |
| LOOCV R2 | 0.76 |
| LOOCV RMSE | 1.245865 |
| LOOCV MAE | 0.999091 |
Summary of SVR model performance.
FIGURE 6
FIGURE 7
Unlike penalized linear models, SVR handles nonlinear dependency. The improvement in cross-validated accuracy shows that the formation of COR includes nonlinear dependency, which is not fully captured using linear shrinkage. Under extreme scarcity, the ability of the kernel to adapt to the data, along with the strict nested validation, allows for better performance without compromising the integrity of the generalization.
4.4 Cross-model comparative analysis and selection
The results of the comparative analysis establish a methodological hierarchy even in an extreme situation of data scarcity. The unregularized linear model confirms that statistical significance and good in-sample fit are not sufficient to guarantee predictive reliability; variance inflation prevails when there are no constraints on complexity. Stability is significantly enhanced by regularization. LASSO shows that only DC has a strong enough linear signal to generalize, although Ridge provides a more balanced distribution of explanatory power among correlated predictors, albeit to a lesser extent in terms of stability gain. This difference in sparsity and smoothness suggests that, in small multicollinear datasets, signal concentration may be more protective against variance-induced instability than smoothing of coefficients. The nonlinear SVR model has the strongest cross-validated performance, as measured by nested validation, indicating that COR formation has curvature and interaction structure in addition to linear effects. Linear shrinkage provides stability, but learning the kernel identifies the dependency patterns, which are significant in terms of fidelity. Collectively, this evidence suggests that for accurate prediction of COR within a fragmented public sector system, strict control of variance is necessary as well as controlled nonlinear flexibility. Accordingly, the SVR specification is chosen as the final model based on the most stable balance of complexity, generality, and structural expressiveness.
5 Model interpretation and governance implications
5.1 Sensitivity structure and nonlinear influence
The response surfaces presented in Figure 8 show that the process of COR development is not based on symmetric proportional-linear relationships, but rather on asymmetrical non-linear relationships with thresholds. DC has the steepest structural gradient, as shown in Figure 8c, where the predicted COR decreases dramatically with increasing maturity. The curve is convex for lower values of DC, and this shows that there is an amplification effect: Incomplete documentation leads to a higher probability of change, but at the same time makes the coordination architecture less stable. Change order volatility behaves as an emergent property of information incompleteness. The nonlinear escalation pattern for RFIs, as shown in Figure 8d, has minimal marginal effects at lower density levels, with increased effects at higher density levels. DUR has a non-monotonic escalation pattern, as shown in Figure 8b, which might indicate exposure modulation effects. TV and CL display weakened marginal effects, while GRL displays a constant monotonic escalation pattern, as shown in Figure 8e, which might indicate technical uncertainty effects. The magnitudes of the total effects in Figure 9 verify the hierarchy of DC, followed by RFIs and DUR. Informational robustness is more significant than financial scale in explaining COR variation. These sensitivity curves and the tornado chart are model-based diagnostics derived from one-variable-at-a-time perturbation of the trained SVR, and they should be interpreted as indicative patterns rather than causal effects.
FIGURE 8
FIGURE 9
5.2 Governance leverage and strategic implications
The sensitivity hierarchy recasts the instability of change orders fundamentally as a failure of information governance rather than a failure of contractual enforcement. This extends prior descriptive findings by quantifying, under strict validation, the dominance of documentation maturity and RFIs over project scale in shaping COR outcomes in the studied context. The strong nonlinear amplification implied by the concept of design completeness suggests that small improvements in documentation maturity are associated with large decreases in volatility. Change orders are thus conceived not as individual breaches of contract, but as a systemic outcome of the informational brittleness introduced at tender stage. Interdisciplinary validation, documentation audits, and digital coordination are not merely implemented as process improvements; rather, they are conceived as a risk dampening structure.
RFIs are found to reveal themselves as instability indicators. Increases in RFI density are found to indicate design ambiguity before financial instability, thereby transforming what was previously considered an administrative metric into a predictive metric of governance. Similarly, the monotonic effect of geotechnical risk level supports the presence of unresolved uncertainty in structure redesign, thereby supporting the importance of validation standards instead of discretionary investigation practices.
At the strategic level, the results show that change order volatility is an emergent attribute of information governance quality, not project magnitude or contractual stringency. From the decision maker’s point of view, public sector resilience is engineered upstream via documentation integrity, technical validation, and information control. There is no downstream administrative tightening that can overcome upstream information weakness. From the public sector point of view, predictive stability is an attribute of the governance architecture, not the intensity of enforcement.
6 Limitations and future research
The main limitation of this study is the extremely small sample size of 21 public projects, which reduces statistical power and increases the sensitivity of the results to individual observations. In addition, all projects originate from a single national public sector context, so the findings may not generalize to other governance systems or market conditions. Some of the key predictors, such as design completeness and complexity level, rely on expert judgment and coded categories, which may introduce subjectivity despite the use of consistent criteria.
Future research should therefore extend this work using larger datasets from multiple agencies and countries and should include external validation on fully independent samples. It would also be useful to refine the measurement of documentation quality and RFIs (for example, through normalized or more granular indicators), to compare alternative machine learning algorithms under the same leakage-free validation pipeline, and to explore how such models can be embedded in practical decision-support tools for public owners.
7 Discussion
This study is motivated by a practical gap in public construction delivery under extreme data scarcity: public owners often experience change-order volatility but lack tender-stage tools that can translate the limited records available into an early, decision-support signal. In many public systems, change orders are still managed reactively, after cost and schedule impacts have already materialized. The purpose of the proposed governance-calibrated pipeline is therefore not only prediction as an academic exercise, but earlier governance actions such as prioritizing design reviews, documentation audits, and investigation depth—before contract award.
A second motivation is that much of the existing change-order literature is descriptive and retrospective. While several drivers of change orders are already discussed in practice and academia, the key question for tender-stage governance is which signals remain robust enough to guide decisions when evaluated on unseen projects in a fragmented public dataset. The comparative modeling results in this study indicate that unconstrained linear relationships can appear convincing within a sample but may not transfer reliably to new projects. This supports the need for disciplined variance control and leakage-free validation so that the study’s conclusions reflect stable patterns rather than chance fit.
The results also refine prior notions about what matters most. Across models and sensitivity analyses, documentation maturity at tender stage emerges as the most consistent signal linked to subsequent COR behavior. This supports a governance interpretation in which change-order volatility is driven primarily by informational weakness introduced upstream. Requests for Information further act as early warning indicators: an increasing density of RFIs reflects ambiguity and coordination gaps that often precede financial instability. In contrast, factors that are commonly assumed to dominate change exposures such as project scale or contractual strictness—show weaker leverage in the presented sensitivity structure when compared with documentation maturity and ambiguity signals.
Finally, the use of a nonlinear learning component is justified by the observed threshold-type behavior in the response surface. The sensitivity patterns suggest that COR does not increase in a smooth proportional manner; instead, limited reductions in design completeness can lead to disproportionate escalation in change exposure, especially when combined with other uncertainty sources. In this sense, the pipeline is designed to capture the practical reality that tender-stage governance variables interact and amplify each other, rather than acting as isolated linear contributors. Overall, the study frames COR predictability as a function of governance quality—especially documentation completeness and ambiguity management, providing a practical basis for upstream interventions to reduce post-award volatility.
8 Conclusion
This study shows that COR prediction is achievable under severe data scarcity conditions if variance control validation is implemented rigorously. The performance of the unregularized model was strong on the training data (R2 of 0.84), though it performed poorly on LOOCV (R2 of 0.54). This supports the argument that statistical significance is not a guarantee of prediction accuracy on small data sets with strong correlations. The performance of the regularized models was better on LOOCV: LASSO improved the LOOCV R2 to 0.70 by identifying design completeness as the sole linear predictor of COR. Ridge regularization improved LOOCV performance to R2 of 0.60 through smooth shrinkage of the weights. The nonlinear SVR model performed the best on LOOCV (R2 of 0.76; RMSE of 1.25), suggesting that COR formation is nonlinear and involves curvature relationships that are not fully captured by linear models. The technical contribution of this study is a leakage-free validation framework that is applicable to data-scarce public sector data environments where variance governance is more important than algorithmic sophistication.
In addition, sensitivity analysis shows that design completeness is nonlinear amplification: increases in design maturity are associated with disproportionate decreases in COR estimates. RFIs are instability leaders, and geotechnical risk is a source of cumulative uncertainty. The results recast change order volatility as a property of informational governance quality rather than project scale or contractual severity.
For decision-making purposes, resilience is pre-emptively achieved through integrity of documentation, technical validation, and ambiguity management; downstream administrative tightening is ineffective as a corrective strategy for upstream informational weakness.
Given the very small sample size (21 projects) and the single-country context, the findings should be interpreted with caution and validated in larger and more diverse datasets.
Statements
Data availability statement
The data supporting the findings of this study are not publicly available due to contractual and confidentiality constraints but are available from the corresponding author upon reasonable request.
Author contributions
MA-B: Methodology, Validation, Writing – original draft. HA: Conceptualization, Software, Writing – review and editing. IA: Data curation, Investigation, Writing – review and editing. JA-B: Formal Analysis, Investigation, Writing – review and editing. TR: Project administration, Supervision, Writing – review and editing. AA: Conceptualization, Formal Analysis, Writing – original draft. Writing – review and editing.
Funding
The author(s) declared that financial support was not received for this work and/or its publication.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that generative AI was not used in the creation of this manuscript.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
References
1
AbuaddousM.Al-BtooshJ. A. A.Al-BtoushM. A. K. A.AlkherretA. J. (2020). Building information modeling strategy in mitigating variation orders in roads projects. Civ. Eng. J.6 (10), 1974–1982. 10.28991/CEJ-2020-03091596
2
Al KailaniH.SweisG. J.SammourF.MaaitahW. O.SweisR. J.AlkailaniM. (2025). Predicting construction cost index using fuzzy logic and machine learning in Jordan. Constr. Innov.25 (5), 1479–1500. 10.1108/CI-08-2023-0182
3
AlmashaqbehH. K.IrshidatM. R.NajjarY. (2022a). Using ANN to predict post-heating mechanical properties of cementitious composites reinforced with multi-scale additives. Smart Struct. Syst.29 (2), 337–350. 10.12989/sss.2022.29.2.337
4
AlmashaqbehH. K.IrshidatM. R.NajjarY.ElmahmoudW. (2022b). Artificial neural network modeling to predict the flexural behavior of RC beams retrofitted with CFRP modified with carbon nanotubes. Comput. Concr.30 (3), 209–224. 10.12989/cac.2022.30.3.209
5
AlsulamyS. (2025). Predicting construction delay risks in Saudi Arabian projects: a comparative analysis of CatBoost, XGBoost, and LGBM. Expert Syst. Appl.268, 126268. 10.1016/J.ESWA.2024.126268
6
AlzubiY.AljaafrehA.KhatatbehA. (2024). Application of machine learning techniques in estimating the construction cost of residential buildings in the Middle East region. Int. J. Constr. Manag.24 (9), 946–958. PAGE:STRING:ARTICLE/CHAPTER. 10.1080/15623599.2023.2239494
7
Aragonés-BeltránP.García-MelónM.Montesinos-ValeraJ. (2017). How to assess stakeholders’ influence in project management? A proposal based on the analytic network process. Int. J. Proj. Manag.35 (3), 451–462. 10.1016/J.IJPROMAN.2017.01.001
8
ArlotS.CelisseA. (2010). A survey of cross-validation procedures for model selection, Stat. Surv.4(none), 40–79. 10.1214/09-SS054
9
AsghariV.Hossein KazemiM.ShahrokhishahrakiM.TangP.AlvanchiA.HsuS. C. (2023). Process-oriented guidelines for systematic improvement of supervised learning research in construction engineering. Adv. Eng. Inf.58, 102215. 10.1016/j.aei.2023.102215
10
AssafS. A.Al-KhalilM.Al-HazmiM. (1995). Causes of delay in large building construction projects. J. Manag. Eng.11 (2), 45–50. 10.1061/(ASCE)0742-597X(1995)11:2(45)
11
AssbeihatJ. M.SweisG. J. (2015). Factors affecting change orders in public construction projects. Int. J. Appl. Sci. Technol.5 (6).
12
AyhanM.DikmenI.BirgonulM. T. (2021). Predicting the occurrence of construction disputes using machine learning techniques. J. Constr. Eng. Manag.147 (4), 04021022. 10.1061/(ASCE)CO.1943-7862.0002027
13
BassamA.Al-BtoushM. A. K. A. (2025). Integrating machine learning and metaheuristic optimization into BIM frameworks for mitigating variation orders: evidence from the Jordanian construction sector. Asian J. Civ. Eng.2025, 1–13. 10.1007/S42107-025-01551-0
14
BatoolI.GoldmannK. (2021). The role of public and private transport infrastructure capital in economic growth. Evidence from Pakistan. Res. Transp. Econ.88, 100886. 10.1016/J.RETREC.2020.100886
15
ChattapadhyayD. B.PuttaJ.Rama Mohan RaoP. (2021). Risk identification, assessments, and prediction for mega construction projects: a risk prediction paradigm based on cross analytical-machine learning model. Buildings11 (4), 172. 10.3390/BUILDINGS11040172
16
ChengJ.DekkersJ. C. M.FernandoR. L. (2021). Cross-validation of best linear unbiased predictions of breeding values using an efficient leave-one-out strategy. J. Animal Breed. Genet.138 (5), 519–527. 10.1111/JBG.12545;WGROUP:STRING:PUBLICATION
17
ChengM. Y.VuQ. T.GosalF. E. (2025). Hybrid deep learning model for accurate cost and schedule estimation in construction projects using sequential and non-sequential data. Automation Constr.170, 105904. 10.1016/J.AUTCON.2024.105904
18
DruckerH.BurgesC. J.KaufmanL.SmolaA.VapnikV. (1996). Support vector regression machines. Adv. Neural Inf. Process. Syst.9.
19
DuK. L.JiangB.LuJ.HuaJ.SwamyM. N. S. (2024). Exploring kernel machines and support vector machines: principles, techniques, and future directions. Mathematics12 (24), 3935. 10.3390/MATH12243935
20
FanC. L. (2022). Evaluation of classification for project features with machine learning algorithms. Symmetry14 (2), 372. 10.3390/SYM14020372
21
FlyvbjergB. (2014). What you should know about megaprojects and why: an overview. Proj. Manag. J.45 (2), 6–19. 10.1002/PMJ.21409
22
GarciaG. R.MichauG.EinsteinH. H.FinkO. (2021). Decision support system for an intelligent operator of utility tunnel boring machines. Automation Constr.131, 103880. 10.1016/j.autcon.2021.103880
23
HamdanM. M.ThneibatM.HyariK. (2026). Predicting cost overrun in construction projects using machine learning algorithms: the case of Jordan. Eng. Constr. Archit. Manag.33 (6), 4732–4763. 10.1108/ECAM-09-2024-1209/1259970
24
HannaA. S.CamlicR.PetersonP. A.NordheimE. V. (2002). Quantitative definition of projects impacted by change orders. J. Constr. Eng. Manag.128 (1), 57–64. 10.1061/(ASCE)0733-9364(2002)128:1(57)
25
HastieT.TibshiraniR.FriedmanJ. (2009). The Elements of Statistical Learning. Springer Series in Statistics. 10.1007/978-0-387-84858-7
26
HiyassatM. A.AlkasagiF.El-MashalehM.SweisG. J. (2022). Risk allocation in public construction projects: the case of Jordan. Int. J. Constr. Manag.22 (8), 1478–1488. 10.1080/15623599.2020.1728605
27
JaberF. K.Al-ZwainyF. M. S.HachemS. W. (2019). Optimizing of predictive performance for construction projects utilizing support vector machine technique. Cogent Eng.6 (1), 1685860. STRING:ARTICLE/CHAPTER. 10.1080/23311916.2019.1685860
28
KoszykowskiM.OrzeszkoW. (2025). Machine learning in project schedule creation: a systematic literature review. J. Sched.2025, 1–18. 10.1007/S10951-025-00857-W
29
LiuH.TianG. (2019). Building engineering safety risk assessment and early warning mechanism construction based on distributed machine learning algorithm. Saf. Sci.120, 764–771. 10.1016/J.SSCI.2019.08.022
30
LumumbaV. W.KiprotichD.Lemasulani MpaineM.Grace MakenaN.Daniel KavitaM. (2024). Comparative analysis of cross-validation techniques: LOOCV, K-folds cross-validation, and repeated K-folds cross-validation in machine learning models. SSRN Electron. J.10.2139/SSRN.5266507
31
ManfrediP. (2026). Polynomial chaos vs kernel machine learning methods for uncertainty quantification: a comparative study and benchmarking with the hybrid PCE-GPR approach. Comput. Methods Appl. Mech. Eng.449, 118523. 10.1016/J.CMA.2025.118523
32
NajiK. K.GunduzM.NaserA. F.NajiK. K.GunduzM.NaserA. F. (2022). Construction change order management project support system utilizing Delphi method. J. Civ. Eng. Manag.28 (7), 564–589. 10.3846/jcem.2022.17203
33
OhJ.TouranA.D’AngeloD.ClarkT.FisherC.GaskinsC.et al (2025). Machine learning–enhanced recurrent event modeling for change order recurrence in highway construction. J. Manag. Eng.41 (5), 04025032. 10.1061/JMENEA.MEENG-6583;WGROUP:STRING:PUBLICATION
34
OttavianiF. M.Ballesteros-PérezP.NarbaevT. (2025). Automated machine learning pipeline for robust project cost and duration forecasting. Automation Constr.178 (1), 106426. 10.1016/j.autcon.2025.106426
35
PerkinsR. A. (2009). Sources of changes in design–build contracts for a governmental owner. J. Constr. Eng. Manag.135 (7), 588–593. 10.1061/(asce)0733-9364(2009)135:7(588)
36
Redha GherabaA.FedlerC. B.PatiD.SchmidtM.DarwishM.NejatA.et al (2023). Causes, effects, and control measures of construction change orders in the U.S. J. Civ. Eng. Archit.17, 57–73. 10.17265/1934-7359/2023.02.001
37
Rudžianskaite-KvaraciejieneR.ApanavičieneR.GelžinisA. (2015). Modelling the effectiveness of PPP road infrastructure projects by applying random forests. J. Civ. Eng. Manag.21 (3), 290–299. PAGE:STRING:ARTICLE/CHAPTER. 10.3846/13923730.2014.971129
38
SantosJ. I.PeredaM.AhedoV.GalánJ. M. (2023). Explainable machine learning for project management control. Comput. and Industrial Eng.180 (1), 109261. 10.1016/j.cie.2023.109261
39
ShihadehJ.Al-ShaibieG.BisharahM.AlshamiD.AlkhadrawiS.Al-BdourH. (2024). Evaluation and prediction of time overruns in Jordanian construction projects using coral reefs optimization and deep learning methods. Asian J. Civ. Eng.25 (3), 2665–2677. 10.1007/S42107-023-00936-3
40
ShugranA. A.GhazaliF. E. M. (2025). Understanding the effects of variation orders on construction project success: insights from the Jordanian context. Int. J. Constr. Manag.25 (5), 519–529. PAGE:STRING:ARTICLE/CHAPTER. 10.1080/15623599.2024.2337238
41
SobaihA. M. H., E. E.ShenW.XueJ.IsmaeilE. M. H.ElnasrA.SobaihE. (2024). A proposed model for variation order management in construction projects. Buildings14 (3), 726. 10.3390/BUILDINGS14030726
42
SuccarB.SherW.WilliamsA. (2012). Measuring BIM performance: five metrics. Archit. Eng. Des. Manag.8 (2), 120–142. 10.1080/17452007.2012.659506;WGROUP:STRING:PUBLICATION
43
Tayefeh HashemiS.EbadatiO. M.KaurH. (2020). Cost estimation and prediction in construction projects: a systematic review on machine learning techniques. SN Appl. Sci.2 (10), 1703. 10.1007/S42452-020-03497-1/FIGURES/11
44
TaylorT. R. B.UddinM.GoodrumP. M.McCoyA.ShanY. (2012). Change orders and lessons learned: knowledge from statistical analyses of engineering change orders on Kentucky highway projects. J. Constr. Eng. Manag.138 (12), 1360–1369. 10.1061/(asce)co.1943-7862.0000550
45
YangX.WenW. (2018). Ridge and lasso regression models for cross-version defect prediction. IEEE Trans. Reliab.67 (3), 885–896. 10.1109/TR.2018.2847353
46
ZengN.LiuY.GongP.HertoghM.KönigM. (2021). Do right PLS and do PLS right: a critical review of the application of PLS-SEM in construction management research. Front. Eng. Manag.8 (3), 356–369. 10.1007/s42524-021-0153-5
47
ZhangW.YuanG.XueR.HanY.TaylorJ. E. (2022). Mitigating common method bias in construction engineering and management research. J. Constr. Eng. Manag.148 (9), 04022089. 10.1061/(asce)co.1943-7862.0002364
48
ZhouH.GaoB.TangS.LiB.WangS. (2025). Intelligent detection on construction project contract missing clauses based on deep learning and NLP. Eng. Constr. Archit. Manag.32 (3), 1546–1580. 10.1108/ECAM-02-2023-0172
Summary
Keywords
change orders, data-scarce construction environments, leave-one-out cross-validation, machine learning, support vector regression, tender stage prediction
Citation
Al-Btoush MAKA, Almashaqbeh HK, Al Shalout I, Aldiabat Al-Btoosh JA, Rawashdeh TM and Ismail Alsmadi A (2026) A governance-calibrated machine learning pipeline for predicting change order ratio under extreme data scarcity in public construction. Front. Built Environ. 12:1837617. doi: 10.3389/fbuil.2026.1837617
Received
24 March 2026
Revised
31 May 2026
Accepted
08 June 2026
Published
08 July 2026
Volume
12 - 2026
Edited by
Pavithra Rathnasiri, Northumbria University, United Kingdom
Reviewed by
Emmanuel Chidiebere Eze, Teesside University, United Kingdom
Vishal Kumar, Massey University, New Zealand
Updates
Copyright
© 2026 Al-Btoush, Almashaqbeh, Al Shalout, Aldiabat Al-Btoosh, Rawashdeh and Ismail Alsmadi.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Mohammed A. KA. Al-Btoush, muhammad.albtoosh@iu.edu.jo
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.