ORIGINAL RESEARCH article

Front. Energy Res., 19 August 2021

Sec. Sustainable Energy Systems

Volume 9 - 2021 | https://doi.org/10.3389/fenrg.2021.707937

A New Two-Stage Approach with Boosting and Model Averaging for Interval-Valued Crude Oil Prices Forecasting in Uncertainty Environments

  • 1. School of Statistics and Mathematics, Central University of Finance & Economics, Beijing, China

  • 2. Academy of Mathematics and Systems Science, Chinese Academy of Sciences, Beijing, China

  • 3. Center for Forecasting Science, Chinese Academy of Sciences, Beijing, China

  • 4. School of Economics and Management, University of Chinese Academy of Sciences, Beijing, China

Abstract

In view of the intrinsic complexity of the oil market, crude oil prices are influenced by numerous factors that make forecasting very difficult. Recognizing this challenge, numerous approaches have been introduced, but little work has been done concerning the interval-valued prices. To capture the underlying characteristics of crude oil price movements, this paper proposes a two-stage forecasting procedure to forecast interval-valued time series, which generalizes point-valued forecasts to incorporate uncertainty and variability. The empirical results show that our proposed approach significantly outperforms all the benchmark models in terms of both forecasting accuracy and robustness analysis. These results can provide references for decision-makers to understand the trends of crude oil prices and improve the efficiency of economic activities.

1 Introduction

As one of the most important commodities, crude oil plays a vital role in various fields. In the past decades, crude oil prices have been extremely volatile (see Figure 1). The oil-related industries are highly sensitive to oil price changes (; ). Accurate prediction of crude oil prices and the market volatility is valuable for market participants to make risk management plans and investment decisions (; ). The crude oil prices are volatile, and are dependent on many factors such as market trends, sentiments and stock markets. The aforementioned factors make the crude oil prices unstable and makes its prediction complicated and challenging. Thus, we aim to develop a reliable model for crude oil price forecasting.

FIGURE 1

In recent literatures, most of the existing methods focus on the point-valued crude oil closing prices (; ; ; ; ; ; ; ). However, the use of closing prices has the disadvantage that it does not take into account the oil price variation information within a given period time, e.g., the midpoint and range of crude oil prices in October 2008 are about /bbl and /bbl respectively. While the midpoint and range of crude oil prices in November 2009 are around /bbl and /bbl respectively.

Such forecasts with point-valued crude oil price data have not been particularly successful when compared with the interval-valued time series forecasts (see ). What is more, recent studies also provide empirical evidence suggesting that ITS models have achieved great success on improving the forecast accuracy in a wide range of fields such as stock price forecasting (; ) and forecasting in energy markets, such as electric power demand (; ), and crude oil prices (). By accessing more information (e.g., highs, lows, midpoints, and range), an interval-based method is expected to be superior to the point-based method (). Here, highs and lows are points of inflection for prices. The price range is the difference between two boundaries, which gives the interval length. It can be regarded as a measure of volatility to reflect the price fluctuation. For example, instead of traditional point-based method, introduce interval dummy variables in the autoregressive conditional interval models. apply a threshold autoregressive interval-valued model. develop an interval-valued factor pricing model. Conclusions from prior studies suggest that interval-valued time series (ITS) models may produce more accurate forecasts.

Therefore, the desirable characteristics of the interval modeling make them ideal candidates for the prediction of crude oil prices. In addition, it is well known that a large set of factors are responsible for changes in the crude oil price, including overall economic conditions, demand and supply, monetary policy, as well as speculative trading (; ). Thus, the number of potential predictors can be very large. In such cases, interval-valued variable selection is considered necessary and becomes the critical step in achieving promising forecasting performances in data-rich environments. On the other hand, in practice, when only some of the variables are selected to include as the predictors in a model, model misspecification is unavoidable, which can worsen the model forecast performance of the model. Therefore, model averaging is considered to take a weighted average of possible combinations of selected interval-valued predictors.

For these reasons, this paper proposes a new two-stage procedure for interval valued crude oil price forecasting based on boosting and model averaging. First, we extend the boosting method by to achieve variable selection for the interval model. Several penalized methods have been proposed to achieve variable selection. Examples include the class of Bridge estimators (), where the Lasso-type estimators are included a special case (), or the smoothly clipped absolute deviation (SCAD) estimator (). Instead of these regularized (penalized) methods, apply information criteria for moment selection, develop boosting for variable selection, where variable selection and shrinkage are performed simultaneously to increase prediction accuracy. The proposed vector boosting algorithm can achieve significant dimension reduction when a long list of interval-valued variables is available.

Next, we extend the LsoMA method developed by to average predictions from interval models with interval-valued exogenous variables to reduce model uncertainty. The idea of model averaging (MA) is first introduced to combine predictions from many forecasting models by and has received great interest in econometrics and statistics. Model averaging is an extension of model selection which can substantially reduce the selection bias induced by selecting only one candidate model. provide a comprehensive summary of previous research on Bayesian model averaging (BMA) where models are weighted by the posterior model probabilities. Unlike BMA, frequentist model averaging (FMA) usually select the optimal weighting with the smallest information criteria scores (; ; ; ; ; ), Mallows model averaging (MMA) by , jackknife model averaging (JMA) by . extend MMA to the situation of the VAR models.

Univariate and bivariate methods are broadly the two main approaches in the interval modeling literature. In the univariate method, models are presented separately for a pair of attributes of interval variables (e.g., midpoint and range). The two attributes are estimated separately (; ), thus only information of one attribute is used in estimating model parameters at a time. Unlike the univariate method, the bivariate method estimates the two attributes simultaneously (e.g., ; ; ; ; ), which is more desirable in ITS forecasting. Therefore, in this paper, in order to consider possible interdependence between midpoint and range, the LsoMA methods are constructed following the bivariate modeling approach to efficiently use the contained information.

This paper proposes a two-stage vector boosting model averaging (2SVBMA) forecasting framework: Stage 1 uses vector Boosting to select interval-valued variables; Stage 2 uses the leave-subject-out cross-validation model averaging method with exogenous interval-valued variables to average interval-valued predictions. Our procedure combines the merits of these two techniques and can be easily adapted to any new situation. We compare our 2SVBMA method with other competing methods including model selection methods by Akaike information criterion (AIC), Bayesian information criterion (BIC), Hannan-Quinn (HQ), and model averaging methods by smoothed AIC, smoothed BIC (), smoothed HQ, and MMA in interval model. The empirical results indicate that the 2SVBMA method has better forecasting performance than the commonly used model selection and averaging methods.

Our proposed 2SVBMA forecasting procedure has a few appealing features. First, this approach extends the forecasting success of point-valued data models of crude oil price to interval-valued data models, which is capable of assessing and forecasting the changes in both the trend and volatility of crude oil prices simultaneously due to the informational gain from interval-valued data. Second, our vector boosting method provides a parsimony and feasible solution to the interval-valued variable selection problem for interval models. Third, the extended interval-valued LsoMA model with interval-valued exogenous variables demonstrates the gains in forecast accuracy through forecast combination. By doing so, our approach improves crude oil price forecasting performances significantly.

The remainder of this paper is organized as follows. Section 2 first proposes 2SVBMA methodology, starts with extended boosting to interval-valued variable selection and develops the LsoMA with interval-valued model with interval-valued exogenous variables. Section 3 provides the empirical implementations. Section 4 discusses the empirical results. Section 5 concludes.

2 Methodology

2.1 Model Framework

Let be a probability space, where Ω is the set of elementary events, is the -field of events, and is the -additive probability measure. An interval random variable is defined as a measurable mapping , such that for all there is a set , where with (; ). A stochastic ITS can be represented by its midpoint and range, i.e., , where and . Assume that {yt} is stationary and follows a vector autoregressive models with interval-valued exogenous variables:where , , and is an interval-valued sequence with mean zero and covariance matrix , and and are the coefficient matrix that satisfies and , is a vector, is a vector, and the assumed initial data are . This data generating process guarantees the natural order of the intervals, i.e., the lower bound is smaller than or equal to the upper bound.

In matrix form, (1) is represented by

andwhere , , , , , and .

The least squares estimators of and are given by

and

2.2 First Stage: Vector Boosting

We first extend Boosting regularization method to interval model to select a subset of interval-valued variables. is the row in . They are the potential interval-valued variables that will be selected by vector boosting. is the element in and Πk is the corresponding interval-valued coefficient of , where . Let denote the iteration in the vector boosting procedure, and denote the maximum number of iteration. At each step , the interval-valued variable that is most relevant to the “current interval-valued residual” is selected. Denote as the strong learner and as the weak learner for . Let , and .

Vector

Boosting performs an interval-valued variable selection for

using the following procedure:

  • 1. When , the initial weak learner for is

  • 2. For each step.

    • 1) Compute the “current interval-valued residual,” .

    • 2) Regress the current interval-valued residual on each . The estimator is obtained as

The interval-valued variables that has the minimum sum of squared residuals is picked up, such that

  • 3) The weak learner is

where

is the interval-valued variable that is selected.

  • 4) The strong learner is updated as

with

, where

is a learning rate, which can be seen as a small step size when updating

.

To avoid overfitting, a version of AIC is used to choose the optimal number of iteration . Define to be an matrix. From Equation (10),The strong learner at each step is

AIC is given aswhere . Then .

2.3 Second Stage: LsoMA

After selecting these important exogenous interval-valued variables, LsoMA technique is extended to interval candidate models with interval-valued exogenous variables, which is adopted to reduce model uncertainty and increase forecast accuracy.

Consider candidate models used to approximate the DGP in Eq. (1) with to be infinite if the sample size is going to infinity. The th () candidate model is given bywhere , , and . Then in matrix form, we havewhere , , and . For each candidate model, we use multivariate least squares (LS) method to estimate parameters and thus the LS estimator of is , and the corresponding estimator of conditional mean is in th candidate model.

Let the weight vector . Then the model averaging estimator of conditional mean is . To obtain the optimal weights, it is common to minimize the following squared loss function:

However, this loss is infeasible because of the unknown conditional mean . We follow the spirit of to use the following feasible leave-subject-out cross-validation criterion of choosing weightswhere , , is the selected matrix to select observations at time point , is the leave-subject-out cross-validation estimator after deleting some observations around , and ; see more discussions in . Minimizing this criterion, we have

and thus the model averaging estimator is . As proved, the weight obtained by minimizing the feasible cross-validation criterion is asymptotically optimal in the sense of achieving the lowest possible quadratic errors, i.e.,

This shows that the squared error loss obtained from the selected weight vector is asymptotically equivalent to the infeasible optimal averaging estimator.

3 Empirical Implementations

This section applies the proposed 2SVBMA procedure to forecast the real price of crude oil. Data and preliminary analysis are introduced in Section 3.1. Then the selected interval-valued factors are introduced in Section 3.2. Section 3.3 introduces the candidate models. Section 3.4 provides competing methods.

3.1 Data and Preliminary Analysis

Following , and , the daily point-valued WTI crude oil prices are used to construct the interval-valued monthly prices. and denote the daily maximum and minimum prices within th month. and are the midpoint and range from an interval-valued price observation . The data period used in the research is from January 2005 to December 2017. Data on crude oil prices are collected from the US Energy Information Administration (EIA). Figure 2 presents the interval-valued crude oil prices: the range (, right y-Axis), the maximum (, left y-Axis), and minimum (yL,t, left y-Axis) prices within 1 month, where we can see that the boundaries and ranges are interlinked, e.g., a strong increase in volatility () is accompanied by a significant decrease in crude oil prices during the second half of 2008.

FIGURE 2

Table 1 presents the summary of statistical characteristics. First, it is shown that the spread of ranges is slightly smaller than the volatility in the boundaries ( and ), where is the monthly prices from EIA. In addition, the skewness and leptokurtic kurtosis are different among , and . Compared with DyL,t and , is with greater skewness and higher leptokurtic. We can see from Table 1 that the interval-valued data can capture more information than the point-valued data.

TABLE 1

MeanMedianMaximumMinimumStd. devSkewnessKurtosis
75.6373.19145.3132.7424.320.35−0.71
67.3165.26122.3026.1922.780.23−1.01
71.4169.54133.8830.3223.540.30−0.86
0.060.060.32−0.150.080.320.93
−0.06−0.040.13−0.640.11−1.956.30
0.120.100.490.040.072.118.33
0.000.010.22−0.390.09−1.032.83

Basic statistical analysis on monthly interval-valued crude oil prices.

3.2 Interval-Valued Control Variables in the First Stage

The potential choices of monthly interval-valued explanatory variables from various aspects are considered in this section, including the stock market, commodity market, technology factor, search query data, speculation, monetary market and currency market (; ; ; ; ); see Table 2 for more discussions. First, the Augmented Dickey-Fuller tests suggest that the null hypothesis for the original control variables is hardly rejected at the 5% significance level, except for non-commercial net long ratio () and the Federal funds rate (). For stationarity, we use the Hukuhara’s difference of interval-valued exogenous variables. The Hukuhara’s difference between a pair of intervals is essentially equal to the regular difference between points in intervals. As mentioned, the concept of interval with Hukuara’s difference is useful and suitable for econometric analysis of interval data. Take S&P 500 index () as an example. It is defined as , where is the Hukuhara’s difference between intervals, and is the regular difference between intervals. This implies that the midpoints and centers of these interval-valued exogenous variables are stationary after Hukuhara’s difference. Similarly, we have , , and ; see specific definitions in Table 2.

TABLE 2

VariablesDescriptionTransformationExplanation
S&P 500 indexAffect expected cash flows and/or discount rates,
Dow Jones industrial indexbe affected through the expected rate of inflation and the expected real interest rate
COMEX gold future closing pricesSafe haven against oil price movements
LME copper future closing prices
WTI-Brent spot price spreadLevelMeasure of the technology influence
Federal funds rateLevelAs oil prices increased, so did concerns about increasing inflation
Generalized real US dollar indexOil price is dollar-denominated
The key word of oil price in the Google trend search engineLevelReflect psychological behaviors of investors
Non-commercial net long ratioLevelProvide liquidity to offset risks

Monthly interval-valued exogenous variables.

Note: (1) These interval-valued variables after transformations are used in candidate models. Transformations are (i) level: ; (2) : ; (iii) : , where is the original series obtained from EIA or Wind database.

Second, Table 3 provides a summary of statistical characteristics. It is shown that no matter whether the time series is transferred by Hukuhara’s difference, the midpoints and ranges for interval-valued control variables appear to have different skewness and leptokurtic kurtosis properties. This suggests that using one attribute of ITS contains partial information only. Thus, it is highly desirable to utilize the information contained in interval-valued data.

TABLE 3

MeanMedianMaximumMinimumStd. devSkewnessKurtosis
0.050.040.310.010.043.6317.41
0.000.000.06−0.160.03−1.827.59
0.050.040.280.010.043.4716.36
0.000.000.05−0.140.03−1.595.90
0.070.060.240.020.031.774.24
−1.81−1.73−1.22−2.580.31−0.55−0.35
0.090.080.510.020.063.0617.02
1.811.732.521.260.320.54−0.49
2.291.6915.360.012.152.399.12
1.120.8612.24−2.871.772.039.76
0.280.210.970.040.191.140.80
3.884.014.542.600.47−0.930.30
0.030.030.120.010.021.482.97
0.110.120.25−0.090.07−0.23−0.70
0.030.020.100.000.020.900.84
−0.15−0.180.01−0.210.061.921.88
0.190.092.750.010.324.6128.40
1.340.285.320.061.771.260.08

Basic statistical analysis on monthly interval-valued explanatory variables.

Third, we use the extended Boosting regularization method to select interval-valued control variables. Specifically, we set the lag length for every control variable and thus the number of the potential explanatory interval-valued variables equals . For vector boosting, we start with the learning rate , iteration = 100 times. These parameters are adjusted during training. After using various training sets, , , , , , and are selected with duplicates removed and used to do h-step-ahead out-of-sample forecasts of interval-valued crude oil prices.

Furthermore, these selected interval-valued control variables have important economic interpretation for crude oil prices as follows:

: It provides information of fundamentals and volatility contained in S&P 500. The movement of S&P 500 Index may closely mirror that of the crude oil prices (e.g., ; ; ; ). As discussed in and , the oil price shocks influence stock prices by affecting expected cash flows and discount rates, since crude oil is an important input in production and its price can influence the costs for the manufacturing and transport sectors.

(j = 1,2,3): It is the logarithmic difference between Comex gold future prices at and , which provides information in Comex gold future market (e.g., ; ; ; ). Gold serves as store of value especially during periods of economic uncertainties. Oil prices can affect levels of inflation (). Gold investment can be used as a hedge against inflation and currency depreciation. It can also be viewed as a safe haven against the stock market turbulence for investors.

: It is WTI-Brent spot price spread, which is the price difference between crude oil and the byproducts refined from it. The crack spread gives the profit margin that a refinery can expect. Thus, a tight spread can be seen as a indicator that refiners may slow production to tighten supply.

: It is the search query data collected from Internet, which has been widely applied as indicator when analyzing the crude oil prices and has been demonstrated to be effective in improving forecasts performance (; ; ; ). The keyword “oil price” is searched in the Google Trend search engine. Search query data is expected to reflect the psychological aspects of investors when they making strategic investment decisions in the crude oil market ().

3.3 Model Averaging in the Second Stage

3.3.1 Candidate Models

We consider 6 lagged dependent variables and 6 exogenous variables selected from vector boosting. As we use monthly interval-valued crude oil prices, the maximum lag is set to 6, including the past half year information. Exogenous variables are sorted by relevance to during the estimation period. Then, 12 nested interval predictive candidate models are considered as:

Model 1. .

Model 2. .

Model 3. .

Model 4. .

Model 5. .

Model 6. .

Next, 6 exogenous variables are added to Model 6 to construct Models 7–12, sorted by relevance to :

Model 7. .

Model 8. .

Model 9. .

Model 10. .

Model 11. .

Model 12. .

These candidate models are used for LsoMA in the second stage. We do -step-ahead prediction with .

3.4 Competing Methods

In this paper, we compare 2SVBMA forecasts with various competing methods, including AIC, BIC, HQ, Mallows model averaging (MMA; ), smoothed AIC (SAIC), smoothed BIC (SBIC) and smoothed Hannan-Quinn (SHQ) based on the same set of candidate models (model 1 - model 12).

The AIC criterion for the th candidate model is , where minimizes and as the residual covariance matrix from the th candidate model. Similarly, BIC and HQ are model selection methods, minimizing the corresponding criteria , , respectively. These three selected candidate models ares used as benchmark models.

Four model averaging (or forecast combination) methods are considered here. MMA proposed by is an extension of Mallows criterion to vector regression models. Specifically, the multivariate Mallow criterion for model averaging takes the following form:where , , and . The Mallows weight vector is defined by:

SAIC, SBIC and SHQ are simple model averaging methods with the weights

and

and

respectively.

4 Empirical Results

This section compares the forecasting performance of the proposed 2SVBMA approach with various competing methods presented in previous studies by using interval-valued crude oil prices. The whole sample from 2005 January to 2017 December are divided into two parts: one is used for parameter estimation, and the other is used for out-of-sample forecasting. Various subsamples for estimation and forecast are used to test prediction accuracy; see Tables 4, 5.

TABLE 4

Estimation: 2005–2010; Forecast:2011–2013
h2SVBMAMMASAICSBICSHQAICBICHQ
1midpoints0.752.091.581.395.595.59
ranges0.702.621.791.604.993.925.15
4midpoints0.322.171.201.093.503.293.60
ranges1.204.712.792.525.426.865.41
8midpoints0.381.551.060.981.711.931.78
ranges1.052.991.981.853.613.683.64
12midpoints0.392.101.201.093.673.173.67
ranges0.653.201.641.486.894.906.89
Estimation: 2006–2011; Forecast:2012–2014
h2SVBMAMMASAICSBICSHQAICBICHQ
1midpoints0.220.660.470.431.210.911.17
ranges0.331.030.730.662.071.611.84
4midpoints0.260.970.640.581.431.501.34
ranges0.410.970.700.641.391.271.35
8midpoints0.220.750.470.411.291.201.23
ranges0.400.800.500.471.571.221.50
12midpoints0.090.450.220.181.360.611.20
ranges0.250.420.290.280.650.560.66
Estimation: 2007–2012; Forecast:2013–2015
h2SVBMAMMASAICSBICSHQAICBICHQ
1midpoints0.070.160.140.130.380.180.34
ranges0.130.220.190.180.600.260.35
4midpoints0.090.390.240.210.580.740.58
ranges0.180.200.210.200.480.250.29
8midpoints0.170.240.240.300.350.36
ranges0.370.390.460.420.441.360.79
12midpoints0.220.270.250.370.260.35
ranges0.490.730.560.550.830.670.97

MSPE () of the recursive prediction for interval-valued crude oil prices (I).

Note: “Estimation” denotes the sample during this period used to estimate parameters, and “Forecast” denotes the sample during this period used to do out-of-sample forecasts. The best forecasts are marked by boldface, and the second best forecasts are marked by underline.

TABLE 5

Estimation: 2008–2013; Forecast:2014–2016
h2SVBMAMMASAICSBICSHQAICBICHQ
1midpoints0.150.300.290.270.610.340.57
ranges0.811.141.031.001.501.031.38
4midpoints0.390.470.520.500.690.650.70
ranges1.111.711.471.412.152.012.12
8midpoints1.380.930.770.872.830.433.29
ranges2.441.981.611.825.631.035.34
12midpoints0.542.451.180.881.055.845.85
ranges0.902.521.671.524.401.964.54
Estimation: 2009–2014; Forecast:2015–2017
h2SVBMAMMASAICSBICSHQAICBICHQ
1midpoints0.220.550.520.400.471.081.03
ranges1.142.372.121.711.964.454.37
4midpoints1.611.471.141.342.460.662.45
ranges2.033.462.692.485.942.756.01
8midpoints1.671.170.871.056.870.354.15
ranges1.522.922.511.912.2510.069.45
12midpoints5.151.691.061.4213.580.4012.92
ranges0.753.932.081.818.941.518.32

MSPE () of the recursive prediction for interval-valued crude oil prices (II).

Note: “Estimation” denotes the sample during this period used to estimate parameters, and “Forecast” denotes the sample during this period used to do out-of-sample forecasts. The best forecasts are marked by boldface, and the second best forecasts are marked by underline.

Tables 4, 5 report the MSPEs of -step-ahead (1,4,8,12) forecasts for the interval-valued crude oil prices using various estimation and forecast samples. First, it is worth noticing that for the horizons of 1, 4, 8 and 12 months, the 2SVBMA method outperforms other competing methods in most cases; out of the 48 cases considered, with respect to RMSFE of midpoints and ranges, it yields the best outcomes 42 times and the second best outcomes 6 times. Intuitively, the proposed 2SVBMA method selects the important factors at the first stage and then give the optimal weights averaging across the 12 nested regression forecasts. Second, 2SVBMA based on LsoMA outperforms various model averaging and model selection methods, including MMA. One possible explanation is that leave-subject-out cross-validation is more suitable for vector autoregressive situations with heteroscedastic and auto-correlated errors. Additionally, as shown in , the approximate unbiasedness of LsoMA and its asymptotic optimality in terms of obtaining the lowest quadratic errors are established. This is why LsoMA outperforms other model averaging methods (i.e., SAIC, SBIC, and SHQ) in the second stage.

Second, the SBIC estimators always produce the second-best forecasts after the 2SVBMA estimator among all model averaging methods, while SAIC achieves higher forecast criteria than other model averaging methods. Similarly, BIC always yields best forecasts among all model selection methods, while the AIC estimator achieves higher MSFE in most cases. This happens because AIC prefers selecting the relatively complicated model, which is inappropriate for out-of-sample forecasting even though it has good in-sample fitting. A simple model may be better for out-of-sample forecasting.

Furthermore, it is shown that at the second stage, model averaging forecasts outperform model selection forecasts in almost 90% of all cases. The significant advantages of model averaging support the argument of that “model uncertainty and instability seriously impair the forecasting ability of individual predictive regression models.”

Overall, the proposed approach using interval-valued data is capable of assessing and forecasting the changes in both level and volatility. We can see from the results that forecasting with model averaging is generally better than obtaining the predictions from just one model (model selection). Since we may choose a very different model when there are small changes in the original data set, which may lead to a big change in the final conclusions, resulting in non-effective decision-making due to the unstable forecasting process. The proposed method is able to help obtain more stable decision-making when a long list of interval-valued predictors is available in a wide range of fields, for example, the daily trading strategy in the finance field.

5 Conclusion

We propose a novel 2SVBMA forecasting procedure to capture the relevant information available in the interval format and the underlying characteristics of crude oil price movements. Vector Boosting in the first stage and LsoMA in the second stage are extended to interval models with interval-valued exogenous variables. Empirical results show that our proposed approach outperforms other competing model averaging and model selection methods in terms of MSFE of midpoints and ranges.

There are some limitations and potential extensions of our study. First, more advanced optimization algorithms for interval-valued variable selection can be proposed in future work. Second, the candidate models with different structures in model averaging methods can further be developed to enhance forecasting. It would also be interesting to develop interval-based machine learning methods to improve forecast accuracy. Furthermore, the proposed methodology in this paper can be extended to the vector autoregressive (VAR) model, which can cover more applications in economics and finance.

In general, 2SVBMA provides a methodological framework for interval-valued data forecasting when there are a large number of potential predictors. For example, this methodology can be used to quantify the impact of COVID-19 pandemic on oil and gas industry. 2SVBMA can also provide implications for the post-COVID recovery management. The accurate prediction of crude oil prices will assist policy makers in understanding issues affecting different oil industry segments, and help governments be better prepared for the recovery.

6 Compliance With Ethical Standards

The authors thank a number of the participants at Symposium on Interval Data Modelling: Theory and Applications (SIDM 2019) in Beijing for their valuable comments and suggestions. This work was partially supported by National Natural Science Foundation of China (Nos. 71973116, 71988101, 72073126, 72091212), and the disciplinary funding of Central University of Finance and Economics. The authors declare no competing interests. This article does not contain any studies with human participants performed by any of the authors.

Statements

Data availability statement

The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author.

Author contributions

All three authors contributed equally to this work and the order of authorship has nothing other than alphabetical significance.

Funding

This work was partially supported by National Natural Science Foundation of China (Nos. 71973116, 71988101, 72073126, 72091212), the funding of Forecasting and Monitoring of COVID-19 in countries along "Belt and Road" and Related Economic Impacts (ANSO-SBA-2020-12), and the disciplinary funding of Central University of Finance and Economics.

Conflict of interest

The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

References

Summary

Keywords

crude oil prices forecasting, forecast combination, interval-valued time series, model averaging, vector L2-boosting

Citation

Huang B, Sun Y and Wang S (2021) A New Two-Stage Approach with Boosting and Model Averaging for Interval-Valued Crude Oil Prices Forecasting in Uncertainty Environments. Front. Energy Res. 9:707937. doi: 10.3389/fenrg.2021.707937

Received

12 May 2021

Accepted

16 July 2021

Published

19 August 2021

Volume

9 - 2021

Edited by

Farhad Taghizadeh-Hesary, Tokai University, Japan

Reviewed by

Ehsan Rasoulinezhad, University of Tehran, Iran

Robina Iram, Jiangsu University, China

Updates

Copyright

*Correspondence: Yuying Sun,

This article was submitted to Sustainable Energy Systems and Policies, a section of the journal Frontiers in Energy Research

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics