METHODS article

Front. Earth Sci., 01 August 2025

Sec. Solid Earth Geophysics

Volume 13 - 2025 | https://doi.org/10.3389/feart.2025.1601363

Data-driven intelligent productivity prediction model for horizontal fracture stimulation

  • QL

    Qian Li 1

  • YS

    Yiyong Sui 1*

  • ML

    Mengying Luo 2

  • BG

    Bin Guan 3

  • LL

    Lu Liu 4

  • YZ

    Yuan Zhao 5

  • 1. School of Petroleum Engineering, China University of Petroleum (East China), Qingdao, China

  • 2. Petroleum Development Center of ShengLi Oilfield, Dongying, China

  • 3. Engineering Technology Department of Daqing Oilfield Limited Company, Daqing, China

  • 4. Exploration and Development Research Institute of Dagang Oilfield Company, Tianjin, China

  • 5. Tianjin Branch of CNPC Logging, Tianjin, China

Abstract

Traditional methods for predicting post-fracturing productivity in horizontal fractures primarily use fracture and formation parameters for calculations. Complex fracture data are difficult to obtain, and these methods do not consider the effects of displacement mechanisms, fracturing techniques, or time factors on post-fracturing productivity. To address the limitations and shortcomings of existing post-fracturing performance prediction methods for horizontal fractures, a horizontal fracture well productivity prediction model was established by combining physical mechanisms with data-driven approaches. First, based on physical mechanisms, factors influencing well productivity were selected from reservoir properties and fracturing operations. Second, relevant characteristic parameters were chosen from geological conditions, production characteristics, and fracturing techniques to perform clustering analysis on fracturing intervals in the data sample. Intervals with similar multidimensional physical features were grouped into the same category. Under the assumption of similar characteristics and mechanisms, correlation analysis was conducted for each fracturing interval category to identify the dominant controlling factors affecting post-fracturing productivity in each reservoir type. Machine learning algorithms were used to establish intelligent models describing the relationships between post-fracturing production enhancement effects, dominant factors, and production time for each reservoir category. Finally, during fracturing design, the optimal productivity prediction model was matched to each interval based on its characteristics to predict post-fracturing productivity. Additionally, the influence patterns of proppant volume on well productivity were comprehensively analyzed to optimize reasonable proppant volumes for different wells and intervals. Field validation showed that the productivity prediction model achieved an average error of 7.06%, providing a basis for horizontal fracture engineering design and achieving cost reduction and efficiency improvement in oilfield development.

1 Introduction

Fracturing of Vertical wells in shallow oil reservoirs has the characteristics of small scale, low cost, and good effectiveness, is one of the most commonly used methods to increase production. In shallow oil reservoirs, the relatively low vertical stress coupled with a high horizontal in-situ stress dominance results in a propensity for hydraulic fractures to propagate horizontally. As the cornerstone of fracturing design optimization, productivity forecasting for hydraulically fractured wells encompasses dual methodologies: stimulation ratio quantification and production profile prediction. Conventional evaluation of horizontal fracture efficacy predominantly relies on analytical solutions for stimulation ratio (SR) computation. Since pioneered the methodology for predicting productivity of vertically fractured wells via stimulation ratio in 1960, numerous scholars worldwide have conducted extensive studies under diverse conditions to investigate the relationships between stimulation ratios and fracture geometry dimensions coupled with conductivity. proposed a methodology for calculating the SR under steady-state flow conditions in 1961. Cui Disheng () developed a comprehensive SR calculation framework incorporating multiple contributing factors: fracture stimulation effects, formation damage mitigation, drainage area configuration, and well placement optimization, thereby advancing a more holistic predictive model for post-fracturing productivity enhancement. established a correlation between Estimated Ultimate Recovery (EUR) and 10 key production-influencing factors for fractured horizontal wells using Random Forest algorithms. analytically derived stimulation ratios for a centrally located fractured well within a circular drainage area under pseudo-steady state flow conditions. developed a COMSOL-based predictive model quantifying stimulation ratios and liquid production rates through systematic analysis of reservoir properties (permeability, porosity), fracture parameters (length, conductivity), and operational variables (flow rate, bottom hole pressure), explicitly accounting for parameter sensitivity and cross-correlation effects. formulated a semi-analytical single-well model for fractured horizontal wells in heterogeneous tight oil reservoirs, leveraging Laplace-space Green’s functions for rectangular domains. This model specifically evaluates fracture penetration ratio and permeability contrast impacts on post-fracturing productivity. implemented a neuro-fuzzy framework combining fuzzy clustering for well typology classification, Key Performance Indicator analysis for dominant factor identification, and Artificial Neural Networks integrating geological (TOC, brittleness index), completion (stage spacing), and stimulation (proppant intensity) parameters to predict early-phase production. applied grey correlation analysis to identify critical productivity drivers—proppant volume, net pay thickness, total injected fluid volume, and stimulation stages—subsequently constructing a multiple linear regression model for initial production forecasting in tight oil horizontal wells. proposed machine learning-driven productivity prediction models for hydraulically fractured horizontal wells by integrating Artificial Neural Networks and Support Vector Machines (SVM). developed predictive frameworks for post-fracturing production in tight sandstone reservoirs under limited historical datasets, employing an ensemble of Elastic Net regression, Decision Trees, and SVM to identify salient parameters from multi-domain fracture-influencing factors. constructed an XGBoost-based productivity forecasting model for 267 horizontal wells in the Sulige Gas Field, leveraging petrophysical features (gas saturation, porosity) and engineering variables (total proppant volume, cluster spacing) as model inputs to quantify fracture-reservoir interactions.

In summary, analytical formulas for post-fracturing performance prediction inherently require fracture parameters as inputs. However, shallow reservoir fracturing operations are typically low-cost and small-scale, making fracture parameter acquisition challenging. Furthermore, with advancements in multi-fracture stimulation technologies, critical variables such as fracture count are absent in conventional analytical models. These formulas yield static productivity estimates, whereas actual post-fracturing production exhibits dynamic temporal decline behavior. Consequently, analytical approaches demonstrate limited applicability in modern fracturing evaluation. Existing data-driven methodologies predominantly adopt well-centric modeling frameworks while neglecting the impacts of diverse displacement mechanisms (e.g., water flooding, polymer flooding, ASP flooding) on stimulated productivity. In shallow reservoirs undergoing various enhanced oil recovery processes, distinct displacement physics–including viscosity modification (polymer), interfacial tension reduction (surfactant), and mobility control (ASP) – differentially influence fracture-reservoir interactions. Therefore, developing displacement mechanism-specific data-driven productivity prediction models for targeted fracturing intervals shows significant potential to enhance fracturing design accuracy and optimize cost-benefit ratios.

2 Methodology

2.1 Main technical principles

Post-fracturing performance varies significantly across different fracturing intervals, stimulation techniques, and displacement mechanisms due to distinct underlying physical mechanisms. The workflow involves three key steps: (as shown in

Figure 1

):

  • 1. Data Collection and preprocessing: Sample data from wells and target fracturing intervals are collected and preprocessed.

  • 2. Fracturing layer clustering: fracturing intervals are grouped into clusters based on multidimensional features, including geological conditions production characteristics, and stimulation parameters Clusters are formed under the assumption that intervals within the same group share identical attributes, mechanisms, and post-fracturing productivity enhancement patterns.

  • 3. Model Development: For each cluster, dominant factors governing horizontal fracture performance are analyzed. Tailored productivity prediction models are then established for horizontal fracture-dominated intervals, achieving higher accuracy and efficiency for specific geological and operational categories.

FIGURE 1

During fracturing design, the target interval is classified using a clustering model based on its characteristic features. The classified interval category is then matched with the optimal productivity prediction model tailored for that specific interval type. By activating the selected model and inputting relevant interval data and stimulation parameters, the post-fracturing production behavior—including production trends and decline patterns—can be predicted for the target interval. As shown in Figure 2.

FIGURE 2

2.2 Data collection and preprocessing

Based on the physical mechanisms of fracturing-induced productivity enhancement, 12 characteristic parameters were selected, encompassing pre-fracturing geological parameters, production parameters, and fracturing operational parameters, along with two additional features: displacement mechanism and fracturing technique type. The displacement mechanisms primarily include water flooding, polymer flooding, and ASP flooding, while fracturing techniques are categorized into conventional fracturing, multi-fracture stimulation, and selective zonal fracturing. A dataset comprising 5,268 fractured well samples was established, as detailed in Table 1. The distribution of sample data across these categories is presented in Table 2. The dataset structured enables physics-informed machine learning while maintaining operational reality constraints.

TABLE 1

Displacement methodFracturing wells / wellFracturing technologyFracturing wells / well
Water flooding1878Conventional fracturing921
Polymer flooding2379Multi-Fracture Fracturing3267
Alkaline-Surfactant-Polymer1011Selective Fracturing1080
Total5268Total5268

Data composition of the data sample set.

TABLE 2

Data rangePresent formation
pressure /MPa
Well spacing /mDepth in the middle
of the oil layer /m
Sandstone thickness /mEffective thickness /mPorosity /%
Minimum value5.211007980.20.120
Maximum values20.48300119219.312.453
Permeability
/μm2
Fracture uncot
/strip
Total fracturing fluid
volume /m3
Proppant
volume /m3
Pre-fracture water
content /%
Daily oil production from
small layer before
fracking /t
Minimum value0.0041324.5820.001
Maximum values1.52111120998.79

Data range of dataset.

For clustering analysis of fracturing intervals, numerical data are required. In addition to standard preprocessing steps such as normalization, the original dataset contains two categorical/textual variables—displacement mechanisms and fracturing techniques—which were converted into numerical representations using dummy variable encoding (). For displacement mechanisms, ASP flooding was selected as the reference category, with water flooding and polymer flooding encoded as “01” and “10,” respectively. For fracturing techniques, selective zonal fracturing served as the reference category, while conventional fracturing and multi-fracture stimulation were encoded as “01” and “10.” The dataset before and after transformation is presented in Table 3.

TABLE 3

Data parametersData categories before conversionPost-conversion data
Displacement methodAlkaline-Surfactant-Polymer00
Polymer flooding01
Water flooding10
Fracturing technologySelective Fracturing00
Conventional fracturing01
Multi-Fracture Fracturing10

Dummy variable coded transformational relationships between the replacement method and the fracturing process.

2.3 Cluster analysis of fractured intervals

The dataset is partitioned into subsets of similar characteristics based on the similarity between different intervals. Each subset contains samples with closely aligned properties, while maintaining distinct differences between subsets. This approach facilitates a comprehensive understanding of key information patterns within the fracturing well data. The K-means clustering algorithm (; ) — an iterative analytical method—provides interpretable results where cluster centroids represent the characteristic attributes of each group. As clustering outcomes critically depend on the selection of cluster numbers (k), determining the optimal k-value is pivotal, particularly when fracturing interval categories lack predefined definitions.

For this analysis, k-values ranging from 2 to 13 were systematically tested on the fracturing well dataset. Clustering performance was evaluated using the silhouette coefficient metric, which quantifies both intra-cluster cohesion (a(i)) and inter-cluster separation (b(i)) for each k-value. The silhouette score s(i) — calculated as:

Serves as a composite evaluation criterion. Higher s(i) values indicate superior clustering configurations, with the k-value yielding the maximum s(i) identified as the optimal classification scheme for multidimensional fracturing interval features.

As shown in Figure 3, the silhouette coefficient initially increases with the number of clusters (k) and gradually plateaus. The maximum silhouette coefficient is achieved at k = 9, indicating optimal clustering results. Therefore, the fracturing intervals in the dataset exhibit the best clustering performance when classified into 9 categories.

FIGURE 3

2.4 Principal controlling factor analysis

The Maximal Information Coefficient (MIC) is a nonparametric statistical method rooted in mutual information theory, designed to quantify association strength between variables, particularly adept at capturing complex linear and nonlinear relationships in high-dimensional data. Its core principle involves dynamically partitioning data grids to compute the maximum mutual information across varying resolutions. Compared to Pearson’s correlation coefficient, which only identifies linear associations, MIC demonstrates significantly enhanced sensitivity to nonlinear patterns such as exponential, periodic, and piecewise relationships, while maintaining robustness against noise and outliers.

In hydraulic fracturing engineering, nonlinear characteristics frequently govern interactions between reservoir parameters (porosity ϕ, permeability k), operational parameters (proppant volume m, fracture count N), and productivity. Certain parameters exhibit threshold effects on productivity enhancement—where exceeding critical values leads to stabilized stimulation effects—while others follow power-law relationships with production outcomes. Traditional regression models struggle to characterize such complexities, whereas MIC enables precise identification of dominant factors through global optimization of variable association patterns.

Based on geomechanical and flow theory, the following multidimensional parameters were analyzed:

Geological Parameters:

  • Porosity (ϕ)

  • Permeability (k)

  • Sandstone thickness (h)

  • Reservoir mid-depth (D)

Operational Parameters:

  • Proppant volume (m)

  • Fracture count (N)

Dynamic Parameters:

  • Current reservoir pressure (P)

  • Well spacing (L)

  • Pre-fracturing water cut (f)

Target Variable:

  • Post-fracturing productivity (Q)

Z-score standardization applied to eliminate dimensional heterogeneity. For each variable pair (Xi, Q), dynamic grid partitioning was performed in 2D space. Mutual information maxima were computed across grid resolutions. Normalized MIC values (0 ≤ MIC ≤1) were derived, characteristics with a MIC value greater than 0.4 can be considered as primary control factors, with results visualized in Figure 4.

FIGURE 4

The results demonstrate that distinct geological characteristics and microscale mechanistic variations across reservoir categories lead to divergent macro-scale dominant factors governing post-fracturing productivity within each fracturing interval type.

2.5 Intelligent prediction modeling development

The training dataset for each category of fractured well productivity is denoted as D = {(Xi, Qi)} (i = 1,2,3, … ,n), where Xi∈Rp represents the multidimensional feature vector determined by dominant governing factors, and Qi∈R corresponds to the post-fracturing productivity. The dataset D was split into training and testing sets in an 8:2 ratio. Machine learning was performed on the training set to develop post-fracturing productivity prediction models for each fracturing interval category using Gradient Boosted Regression Trees (GBRT), Random Forests (RF), and Bagging. The post-fracturing productivity prediction models were optimized after comparing and analyzing the evaluation metrics.

2.5.1 Random forest models for post-fracturing capacity prediction

The Random forest algorithm is used to establish a post-fracturing capacity prediction model, and the schematic diagram of the Random forest regression algorithm is shown in Figure 5. The maximum values of model tree depth, number of trees and corresponding R2 coefficients of determination for the random forest model for capacity prediction are detailed in Table 4. The average value of the R2 coefficient of determination of the nine types of fracturing well production capacity prediction models established by the random forest regression algorithm is 0.85 for the test set and 0.97 for the training set, which is a difference of 0.12.

FIGURE 5

TABLE 4

Model categories012345678
Regression tree depth/layer151918191716151916
Number of regression trees/tree850100850950150350150900100
Test set coefficient of determinationR20.900.820.830.850.890.840.830.860.82
Training set coefficient of determinationR20.970.960.990.970.980.980.990.960.97

Tree depth, tree number and maximum R2 determination coefficient of 9 types of fracturing well productivity prediction models (Random Forest).

2.5.2 Bagging models for post-fracturing capacity prediction

The Bagging algorithm is used to establish a post-fracturing capacity prediction model, and the schematic diagram of the Bagging regression algorithm is shown in Figure 6.

FIGURE 6

The maximum values of model tree depth, number of trees and corresponding R2 coefficients of determination for the Bagging model for capacity prediction are detailed in Table 5. As shown in Table 5, the R2 coefficient of determination of the nine types of fracturing well production capacity prediction models established by Bagging regression algorithm is 0.87 on average for the test set and 0.95 on average for the training set, with an average difference of 0.09 between the two.

TABLE 5

Model categories012345678
Regression tree depth/layer191920201820191920
Number of regression trees/tree2001008001000200100350600750
Test set coefficient of determinationR20.900.910.850.890.910.850.850.880.82
Training set coefficient of determinationR20.980.940.970.960.980.950.960.950.94

Tree depth, tree number and maximum R2 determination coefficient of 9 types of fracturing well productivity prediction models (Bagging).

2.5.3 GBRT models for post-fracturing capacity prediction

The GBRT algorithm is used to establish a post-fracturing capacity prediction model, and the schematic diagram of the GBRT regression algorithm is shown in Figure 7. During the construction of the productivity prediction models, key parameters considered included tree depth and number of trees. A Bayesian optimization approach was employed to determine the hyperparameters of the nine GBRT models for different fracturing interval categories, as detailed in Table 6.

FIGURE 7

TABLE 6

Model categories012345678
Regression tree depth/layer78107614759
Number of regression trees/tree8009509001000850950850950550
Test set coefficient of determinationR20.910.900.890.890.910.890.900.890.89
Training set coefficient of determinationR20.930.920.930.920.920.910.940.930.92

Maximum values of model tree depth, number of trees, and corresponding R2 coefficients of determination for 9 types of fractured well production capacity prediction models (GBRT).

Comparative analysis of the prediction effect of three 9-class small-layer fracturing well capacity prediction regression models. It can be seen that the Random Forest regression algorithm established by the nine categories of fracturing well capacity prediction model has the smallest R2 coefficient of determination, and the fracturing well capacity prediction GBRT regression model is better than the Bagging model as a whole.

3 Case study

3.1 Proppant volume optimization on the fracturing production rate

The proppant volume is one of the critical parameters influencing fractured well productivity. Designing an appropriate proppant volume prior to fracturing operations not only maximizes the oil-enhancement effects of stimulation but also effectively controls single-well operational costs and improves the cost-benefit ratio (). Insufficient proppant volume leads to inadequate fracture width and uneven proppant distribution, reducing the effective stimulated reservoir volume and reservoir permeability. Conversely, excessive proppant volume may cause proppant flowback or poor packing, hindering the formation of high-conductivity fractures and even resulting in fracturing failure. Therefore, optimizing proppant volume is essential for ensuring fracturing efficacy and enhancing hydrocarbon productivity.

The impact of proppant volume on post-fracturing productivity can be analyzed using predictive models. Sensitivity analysis was conducted on proppant volume for three fracturing intervals from three wells using the productivity prediction model, as shown in Figure 8. Key findings include that proppant volume exhibits distinct impact patterns on productivity across different intervals, yet each interval possesses an optimal proppant volume range. Within this range, fracturing achieves peak production enhancement for the specific interval.

FIGURE 8

3.2 Capacity prediction after hydraulic fracturing

A total of 72 oil wells were randomly selected from out-of-sample datasets to validate the post-fracturing productivity GBRT model. The model predicted production rates for the first month after stimulation, which were then compared with actual field data. As illustrated in Figure 9, the dashed lines represent actual daily incremental oil production, while the solid lines denote predicted values.

FIGURE 9

The performance metrics across the nine model categories in Table 7.

TABLE 7

Model categoriesMRERMSE
07.94%0.26
111.33%0.35
217.12%0.42
314.66%0.45
410.97%0.44
59.92%0.57
612.32%0.58
713.18%0.60
89.47%0.85

Accuracy of prediction models in practical applications.

MRE, ranges from7.94% (Category 0) to 17.12% (Category 2), with an average MRE, of 12.09%. RMSE, values are tightly clustered between 0.26 (Category 0) and 0.85 (Category 8), indicating robust model stability. The models demonstrate strong alignment with field observations, with 6 out of 9 categories achieving MRE <13% and RMSE <0.6, validating their predictive accuracy for post-fracturing performance.

4 Conclusion

  • (1) Based on the physical mechanisms of horizontal fracture fracturing, relevant characteristic parameters were selected from geological conditions, production characteristics, and fracturing techniques to perform clustering analysis on fracturing intervals in the data sample. Similar intervals were categorized, and categorical modeling studies were conducted. This approach allows for matching the best model to different fracturing intervals according to their corresponding categories, thereby improving model applicability.

  • (2) For each category of fracturing intervals, correlation analysis was performed to identify the dominant controlling factors influencing post-fracturing productivity in each reservoir type. The dominant factors affecting post-fracturing performance differ slightly across interval categories. Machine learning algorithms such as Random forest, Bagging and GBRT were used to establish models describing the relationships between post-fracturing production enhancement effects, dominant factors, and production time for each reservoir category. The fracturing well capacity GBRT prediction model can predict productivity after fracturing in different intervals.

  • (3) In the era of smart oilfields, data-driven models hold broad application prospects for horizontal fracture fracturing productivity prediction. They can fully utilize data assets, uncover production patterns, compensate for the limitations of physical models, and improve the accuracy and reliability of post-fracturing productivity predictions.

Statements

Data availability statement

The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.

Author contributions

QL: Methodology, Writing – original draft. YS: Methodology, Writing – review and editing, Formal Analysis. ML: Writing – original draft, Data curation, Software. BG: Methodology, Writing – review and editing. LL: Investigation, Conceptualization, Writing – review and editing. YZ: Conceptualization, Writing – review and editing.

Funding

The author(s) declare that no financial support was received for the research and/or publication of this article.

Conflict of interest

Author BG was employed by Engineering Technology Department of Daqing Oilfield Limited Company. Author LL was employed by Exploration and Development Research Institute of Dagang Oilfield Company. Author YZ was employed by Tianjin Branch of CNPC Logging.

The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

The reviewer FZ declared a shared affiliation with the authors QL and YYS to the handling editor at the time of review.

Generative AI statement

The author(s) declare that no Generative AI was used in the creation of this manuscript.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/feart.2025.1601363/full#supplementary-material

References

Summary

Keywords

data-driven, horizontal fracture fracturing, productivity optimization, application, controlling factors

Citation

Li Q, Sui Y, Luo M, Guan B, Liu L and Zhao Y (2025) Data-driven intelligent productivity prediction model for horizontal fracture stimulation. Front. Earth Sci. 13:1601363. doi: 10.3389/feart.2025.1601363

Received

27 March 2025

Accepted

30 June 2025

Published

01 August 2025

Volume

13 - 2025

Edited by

Xin Sun, Sinopec Matrix Co., LTD., China

Reviewed by

Fuqiong Huang, China Earthquake Networks Center, China

Yingkun Fu, University of Alberta, Canada

Fengjiao Zhang, China University of Petroleum, China

Updates

Copyright

*Correspondence: Yiyong Sui, ,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics