Abstract
In view of the intrinsic complexity of the oil market, crude oil prices are influenced by numerous factors that make forecasting very difficult. Recognizing this challenge, numerous approaches have been introduced, but little work has been done concerning the interval-valued prices. To capture the underlying characteristics of crude oil price movements, this paper proposes a two-stage forecasting procedure to forecast interval-valued time series, which generalizes point-valued forecasts to incorporate uncertainty and variability. The empirical results show that our proposed approach significantly outperforms all the benchmark models in terms of both forecasting accuracy and robustness analysis. These results can provide references for decision-makers to understand the trends of crude oil prices and improve the efficiency of economic activities.
1 Introduction
As one of the most important commodities, crude oil plays a vital role in various fields. In the past decades, crude oil prices have been extremely volatile (see Figure 1). The oil-related industries are highly sensitive to oil price changes (; ). Accurate prediction of crude oil prices and the market volatility is valuable for market participants to make risk management plans and investment decisions (; ). The crude oil prices are volatile, and are dependent on many factors such as market trends, sentiments and stock markets. The aforementioned factors make the crude oil prices unstable and makes its prediction complicated and challenging. Thus, we aim to develop a reliable model for crude oil price forecasting.
FIGURE 1
In recent literatures, most of the existing methods focus on the point-valued crude oil closing prices (; ; ; ; ; ; ; ). However, the use of closing prices has the disadvantage that it does not take into account the oil price variation information within a given period time, e.g., the midpoint and range of crude oil prices in October 2008 are about /bbl and /bbl respectively. While the midpoint and range of crude oil prices in November 2009 are around /bbl and /bbl respectively.
Such forecasts with point-valued crude oil price data have not been particularly successful when compared with the interval-valued time series forecasts (see ). What is more, recent studies also provide empirical evidence suggesting that ITS models have achieved great success on improving the forecast accuracy in a wide range of fields such as stock price forecasting (; ) and forecasting in energy markets, such as electric power demand (; ), and crude oil prices (). By accessing more information (e.g., highs, lows, midpoints, and range), an interval-based method is expected to be superior to the point-based method (). Here, highs and lows are points of inflection for prices. The price range is the difference between two boundaries, which gives the interval length. It can be regarded as a measure of volatility to reflect the price fluctuation. For example, instead of traditional point-based method, introduce interval dummy variables in the autoregressive conditional interval models. apply a threshold autoregressive interval-valued model. develop an interval-valued factor pricing model. Conclusions from prior studies suggest that interval-valued time series (ITS) models may produce more accurate forecasts.
Therefore, the desirable characteristics of the interval modeling make them ideal candidates for the prediction of crude oil prices. In addition, it is well known that a large set of factors are responsible for changes in the crude oil price, including overall economic conditions, demand and supply, monetary policy, as well as speculative trading (; ). Thus, the number of potential predictors can be very large. In such cases, interval-valued variable selection is considered necessary and becomes the critical step in achieving promising forecasting performances in data-rich environments. On the other hand, in practice, when only some of the variables are selected to include as the predictors in a model, model misspecification is unavoidable, which can worsen the model forecast performance of the model. Therefore, model averaging is considered to take a weighted average of possible combinations of selected interval-valued predictors.
For these reasons, this paper proposes a new two-stage procedure for interval valued crude oil price forecasting based on boosting and model averaging. First, we extend the boosting method by to achieve variable selection for the interval model. Several penalized methods have been proposed to achieve variable selection. Examples include the class of Bridge estimators (), where the Lasso-type estimators are included a special case (), or the smoothly clipped absolute deviation (SCAD) estimator (). Instead of these regularized (penalized) methods, apply information criteria for moment selection, develop boosting for variable selection, where variable selection and shrinkage are performed simultaneously to increase prediction accuracy. The proposed vector boosting algorithm can achieve significant dimension reduction when a long list of interval-valued variables is available.
Next, we extend the LsoMA method developed by to average predictions from interval models with interval-valued exogenous variables to reduce model uncertainty. The idea of model averaging (MA) is first introduced to combine predictions from many forecasting models by and has received great interest in econometrics and statistics. Model averaging is an extension of model selection which can substantially reduce the selection bias induced by selecting only one candidate model. provide a comprehensive summary of previous research on Bayesian model averaging (BMA) where models are weighted by the posterior model probabilities. Unlike BMA, frequentist model averaging (FMA) usually select the optimal weighting with the smallest information criteria scores (; ; ; ; ; ), Mallows model averaging (MMA) by , jackknife model averaging (JMA) by . extend MMA to the situation of the VAR models.
Univariate and bivariate methods are broadly the two main approaches in the interval modeling literature. In the univariate method, models are presented separately for a pair of attributes of interval variables (e.g., midpoint and range). The two attributes are estimated separately (; ), thus only information of one attribute is used in estimating model parameters at a time. Unlike the univariate method, the bivariate method estimates the two attributes simultaneously (e.g., ; ; ; ; ), which is more desirable in ITS forecasting. Therefore, in this paper, in order to consider possible interdependence between midpoint and range, the LsoMA methods are constructed following the bivariate modeling approach to efficiently use the contained information.
This paper proposes a two-stage vector boosting model averaging (2SVBMA) forecasting framework: Stage 1 uses vector Boosting to select interval-valued variables; Stage 2 uses the leave-subject-out cross-validation model averaging method with exogenous interval-valued variables to average interval-valued predictions. Our procedure combines the merits of these two techniques and can be easily adapted to any new situation. We compare our 2SVBMA method with other competing methods including model selection methods by Akaike information criterion (AIC), Bayesian information criterion (BIC), Hannan-Quinn (HQ), and model averaging methods by smoothed AIC, smoothed BIC (), smoothed HQ, and MMA in interval model. The empirical results indicate that the 2SVBMA method has better forecasting performance than the commonly used model selection and averaging methods.
Our proposed 2SVBMA forecasting procedure has a few appealing features. First, this approach extends the forecasting success of point-valued data models of crude oil price to interval-valued data models, which is capable of assessing and forecasting the changes in both the trend and volatility of crude oil prices simultaneously due to the informational gain from interval-valued data. Second, our vector boosting method provides a parsimony and feasible solution to the interval-valued variable selection problem for interval models. Third, the extended interval-valued LsoMA model with interval-valued exogenous variables demonstrates the gains in forecast accuracy through forecast combination. By doing so, our approach improves crude oil price forecasting performances significantly.
The remainder of this paper is organized as follows. Section 2 first proposes 2SVBMA methodology, starts with extended boosting to interval-valued variable selection and develops the LsoMA with interval-valued model with interval-valued exogenous variables. Section 3 provides the empirical implementations. Section 4 discusses the empirical results. Section 5 concludes.
2 Methodology
2.1 Model Framework
Let be a probability space, where Ω is the set of elementary events, is the -field of events, and is the -additive probability measure. An interval random variable is defined as a measurable mapping , such that for all there is a set , where with (; ). A stochastic ITS can be represented by its midpoint and range, i.e., , where and . Assume that {yt} is stationary and follows a vector autoregressive models with interval-valued exogenous variables:where , , and is an interval-valued sequence with mean zero and covariance matrix , and and are the coefficient matrix that satisfies and , is a vector, is a vector, and the assumed initial data are . This data generating process guarantees the natural order of the intervals, i.e., the lower bound is smaller than or equal to the upper bound.
In matrix form, (1) is represented by
andwhere , , , , , and .
The least squares estimators of and are given by
and
2.2 First Stage: Vector Boosting
We first extend Boosting regularization method to interval model to select a subset of interval-valued variables. is the row in . They are the potential interval-valued variables that will be selected by vector boosting. is the element in and Πk is the corresponding interval-valued coefficient of , where . Let denote the iteration in the vector boosting procedure, and denote the maximum number of iteration. At each step , the interval-valued variable that is most relevant to the “current interval-valued residual” is selected. Denote as the strong learner and as the weak learner for . Let , and .
Vector
Boosting performs an interval-valued variable selection for
using the following procedure:
1. When , the initial weak learner for is
2. For each step.
1) Compute the “current interval-valued residual,” .
2) Regress the current interval-valued residual on each . The estimator is obtained as
The interval-valued variables that has the minimum sum of squared residuals is picked up, such that
3) The weak learner is
where
is the interval-valued variable that is selected.
4) The strong learner is updated as
with
, where
is a learning rate, which can be seen as a small step size when updating
.
To avoid overfitting, a version of AIC is used to choose the optimal number of iteration . Define to be an matrix. From Equation (10),The strong learner at each step is
AIC is given aswhere . Then .
2.3 Second Stage: LsoMA
After selecting these important exogenous interval-valued variables, LsoMA technique is extended to interval candidate models with interval-valued exogenous variables, which is adopted to reduce model uncertainty and increase forecast accuracy.
Consider candidate models used to approximate the DGP in Eq. (1) with to be infinite if the sample size is going to infinity. The th () candidate model is given bywhere , , and . Then in matrix form, we havewhere , , and . For each candidate model, we use multivariate least squares (LS) method to estimate parameters and thus the LS estimator of is , and the corresponding estimator of conditional mean is in th candidate model.
Let the weight vector . Then the model averaging estimator of conditional mean is . To obtain the optimal weights, it is common to minimize the following squared loss function:
However, this loss is infeasible because of the unknown conditional mean . We follow the spirit of to use the following feasible leave-subject-out cross-validation criterion of choosing weightswhere , , is the selected matrix to select observations at time point , is the leave-subject-out cross-validation estimator after deleting some observations around , and ; see more discussions in . Minimizing this criterion, we have
and thus the model averaging estimator is . As proved, the weight obtained by minimizing the feasible cross-validation criterion is asymptotically optimal in the sense of achieving the lowest possible quadratic errors, i.e.,
This shows that the squared error loss obtained from the selected weight vector is asymptotically equivalent to the infeasible optimal averaging estimator.
3 Empirical Implementations
This section applies the proposed 2SVBMA procedure to forecast the real price of crude oil. Data and preliminary analysis are introduced in Section 3.1. Then the selected interval-valued factors are introduced in Section 3.2. Section 3.3 introduces the candidate models. Section 3.4 provides competing methods.
3.1 Data and Preliminary Analysis
Following , and , the daily point-valued WTI crude oil prices are used to construct the interval-valued monthly prices. and denote the daily maximum and minimum prices within th month. and are the midpoint and range from an interval-valued price observation . The data period used in the research is from January 2005 to December 2017. Data on crude oil prices are collected from the US Energy Information Administration (EIA). Figure 2 presents the interval-valued crude oil prices: the range (, right y-Axis), the maximum (, left y-Axis), and minimum (yL,t, left y-Axis) prices within 1 month, where we can see that the boundaries and ranges are interlinked, e.g., a strong increase in volatility () is accompanied by a significant decrease in crude oil prices during the second half of 2008.
FIGURE 2
Table 1 presents the summary of statistical characteristics. First, it is shown that the spread of ranges is slightly smaller than the volatility in the boundaries ( and ), where is the monthly prices from EIA. In addition, the skewness and leptokurtic kurtosis are different among , and . Compared with DyL,t and , is with greater skewness and higher leptokurtic. We can see from Table 1 that the interval-valued data can capture more information than the point-valued data.
TABLE 1
| Mean | Median | Maximum | Minimum | Std. dev | Skewness | Kurtosis | |
|---|---|---|---|---|---|---|---|
| 75.63 | 73.19 | 145.31 | 32.74 | 24.32 | 0.35 | −0.71 | |
| 67.31 | 65.26 | 122.30 | 26.19 | 22.78 | 0.23 | −1.01 | |
| 71.41 | 69.54 | 133.88 | 30.32 | 23.54 | 0.30 | −0.86 | |
| 0.06 | 0.06 | 0.32 | −0.15 | 0.08 | 0.32 | 0.93 | |
| −0.06 | −0.04 | 0.13 | −0.64 | 0.11 | −1.95 | 6.30 | |
| 0.12 | 0.10 | 0.49 | 0.04 | 0.07 | 2.11 | 8.33 | |
| 0.00 | 0.01 | 0.22 | −0.39 | 0.09 | −1.03 | 2.83 |
Basic statistical analysis on monthly interval-valued crude oil prices.
3.2 Interval-Valued Control Variables in the First Stage
The potential choices of monthly interval-valued explanatory variables from various aspects are considered in this section, including the stock market, commodity market, technology factor, search query data, speculation, monetary market and currency market (; ; ; ; ); see Table 2 for more discussions. First, the Augmented Dickey-Fuller tests suggest that the null hypothesis for the original control variables is hardly rejected at the 5% significance level, except for non-commercial net long ratio () and the Federal funds rate (). For stationarity, we use the Hukuhara’s difference of interval-valued exogenous variables. The Hukuhara’s difference between a pair of intervals is essentially equal to the regular difference between points in intervals. As mentioned, the concept of interval with Hukuara’s difference is useful and suitable for econometric analysis of interval data. Take S&P 500 index () as an example. It is defined as , where is the Hukuhara’s difference between intervals, and is the regular difference between intervals. This implies that the midpoints and centers of these interval-valued exogenous variables are stationary after Hukuhara’s difference. Similarly, we have , , and ; see specific definitions in Table 2.
TABLE 2
| Variables | Description | Transformation | Explanation |
|---|---|---|---|
| S&P 500 index | Affect expected cash flows and/or discount rates, | ||
| Dow Jones industrial index | be affected through the expected rate of inflation and the expected real interest rate | ||
| COMEX gold future closing prices | Safe haven against oil price movements | ||
| LME copper future closing prices | |||
| WTI-Brent spot price spread | Level | Measure of the technology influence | |
| Federal funds rate | Level | As oil prices increased, so did concerns about increasing inflation | |
| Generalized real US dollar index | Oil price is dollar-denominated | ||
| The key word of oil price in the Google trend search engine | Level | Reflect psychological behaviors of investors | |
| Non-commercial net long ratio | Level | Provide liquidity to offset risks |
Monthly interval-valued exogenous variables.
Note: (1) These interval-valued variables after transformations are used in candidate models. Transformations are (i) level: ; (2) : ; (iii) : , where is the original series obtained from EIA or Wind database.
Second, Table 3 provides a summary of statistical characteristics. It is shown that no matter whether the time series is transferred by Hukuhara’s difference, the midpoints and ranges for interval-valued control variables appear to have different skewness and leptokurtic kurtosis properties. This suggests that using one attribute of ITS contains partial information only. Thus, it is highly desirable to utilize the information contained in interval-valued data.
TABLE 3
| Mean | Median | Maximum | Minimum | Std. dev | Skewness | Kurtosis | |
|---|---|---|---|---|---|---|---|
| 0.05 | 0.04 | 0.31 | 0.01 | 0.04 | 3.63 | 17.41 | |
| 0.00 | 0.00 | 0.06 | −0.16 | 0.03 | −1.82 | 7.59 | |
| 0.05 | 0.04 | 0.28 | 0.01 | 0.04 | 3.47 | 16.36 | |
| 0.00 | 0.00 | 0.05 | −0.14 | 0.03 | −1.59 | 5.90 | |
| 0.07 | 0.06 | 0.24 | 0.02 | 0.03 | 1.77 | 4.24 | |
| −1.81 | −1.73 | −1.22 | −2.58 | 0.31 | −0.55 | −0.35 | |
| 0.09 | 0.08 | 0.51 | 0.02 | 0.06 | 3.06 | 17.02 | |
| 1.81 | 1.73 | 2.52 | 1.26 | 0.32 | 0.54 | −0.49 | |
| 2.29 | 1.69 | 15.36 | 0.01 | 2.15 | 2.39 | 9.12 | |
| 1.12 | 0.86 | 12.24 | −2.87 | 1.77 | 2.03 | 9.76 | |
| 0.28 | 0.21 | 0.97 | 0.04 | 0.19 | 1.14 | 0.80 | |
| 3.88 | 4.01 | 4.54 | 2.60 | 0.47 | −0.93 | 0.30 | |
| 0.03 | 0.03 | 0.12 | 0.01 | 0.02 | 1.48 | 2.97 | |
| 0.11 | 0.12 | 0.25 | −0.09 | 0.07 | −0.23 | −0.70 | |
| 0.03 | 0.02 | 0.10 | 0.00 | 0.02 | 0.90 | 0.84 | |
| −0.15 | −0.18 | 0.01 | −0.21 | 0.06 | 1.92 | 1.88 | |
| 0.19 | 0.09 | 2.75 | 0.01 | 0.32 | 4.61 | 28.40 | |
| 1.34 | 0.28 | 5.32 | 0.06 | 1.77 | 1.26 | 0.08 |
Basic statistical analysis on monthly interval-valued explanatory variables.
Third, we use the extended Boosting regularization method to select interval-valued control variables. Specifically, we set the lag length for every control variable and thus the number of the potential explanatory interval-valued variables equals . For vector boosting, we start with the learning rate , iteration = 100 times. These parameters are adjusted during training. After using various training sets, , , , , , and are selected with duplicates removed and used to do h-step-ahead out-of-sample forecasts of interval-valued crude oil prices.
Furthermore, these selected interval-valued control variables have important economic interpretation for crude oil prices as follows:
: It provides information of fundamentals and volatility contained in S&P 500. The movement of S&P 500 Index may closely mirror that of the crude oil prices (e.g., ; ; ; ). As discussed in and , the oil price shocks influence stock prices by affecting expected cash flows and discount rates, since crude oil is an important input in production and its price can influence the costs for the manufacturing and transport sectors.
(j = 1,2,3): It is the logarithmic difference between Comex gold future prices at and , which provides information in Comex gold future market (e.g., ; ; ; ). Gold serves as store of value especially during periods of economic uncertainties. Oil prices can affect levels of inflation (). Gold investment can be used as a hedge against inflation and currency depreciation. It can also be viewed as a safe haven against the stock market turbulence for investors.
: It is WTI-Brent spot price spread, which is the price difference between crude oil and the byproducts refined from it. The crack spread gives the profit margin that a refinery can expect. Thus, a tight spread can be seen as a indicator that refiners may slow production to tighten supply.
: It is the search query data collected from Internet, which has been widely applied as indicator when analyzing the crude oil prices and has been demonstrated to be effective in improving forecasts performance (; ; ; ). The keyword “oil price” is searched in the Google Trend search engine. Search query data is expected to reflect the psychological aspects of investors when they making strategic investment decisions in the crude oil market ().
3.3 Model Averaging in the Second Stage
3.3.1 Candidate Models
We consider 6 lagged dependent variables and 6 exogenous variables selected from vector boosting. As we use monthly interval-valued crude oil prices, the maximum lag is set to 6, including the past half year information. Exogenous variables are sorted by relevance to during the estimation period. Then, 12 nested interval predictive candidate models are considered as:
Model 1. .
Model 2. .
Model 3. .
Model 4. .
Model 5. .
Model 6. .
Next, 6 exogenous variables are added to Model 6 to construct Models 7–12, sorted by relevance to :
Model 7. .
Model 8. .
Model 9. .
Model 10. .
Model 11. .
Model 12. .
These candidate models are used for LsoMA in the second stage. We do -step-ahead prediction with .
3.4 Competing Methods
In this paper, we compare 2SVBMA forecasts with various competing methods, including AIC, BIC, HQ, Mallows model averaging (MMA; ), smoothed AIC (SAIC), smoothed BIC (SBIC) and smoothed Hannan-Quinn (SHQ) based on the same set of candidate models (model 1 - model 12).
The AIC criterion for the th candidate model is , where minimizes and as the residual covariance matrix from the th candidate model. Similarly, BIC and HQ are model selection methods, minimizing the corresponding criteria , , respectively. These three selected candidate models ares used as benchmark models.
Four model averaging (or forecast combination) methods are considered here. MMA proposed by is an extension of Mallows criterion to vector regression models. Specifically, the multivariate Mallow criterion for model averaging takes the following form:where , , and . The Mallows weight vector is defined by:
SAIC, SBIC and SHQ are simple model averaging methods with the weights
and
and
respectively.
4 Empirical Results
This section compares the forecasting performance of the proposed 2SVBMA approach with various competing methods presented in previous studies by using interval-valued crude oil prices. The whole sample from 2005 January to 2017 December are divided into two parts: one is used for parameter estimation, and the other is used for out-of-sample forecasting. Various subsamples for estimation and forecast are used to test prediction accuracy; see Tables 4, 5.
TABLE 4
| Estimation: 2005–2010; Forecast:2011–2013 | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| h | 2SVBMA | MMA | SAIC | SBIC | SHQ | AIC | BIC | HQ | |
| 1 | midpoints | 0.75 | 2.09 | 1.58 | 1.39 | 5.59 | 5.59 | ||
| ranges | 0.70 | 2.62 | 1.79 | 1.60 | 4.99 | 3.92 | 5.15 | ||
| 4 | midpoints | 0.32 | 2.17 | 1.20 | 1.09 | 3.50 | 3.29 | 3.60 | |
| ranges | 1.20 | 4.71 | 2.79 | 2.52 | 5.42 | 6.86 | 5.41 | ||
| 8 | midpoints | 0.38 | 1.55 | 1.06 | 0.98 | 1.71 | 1.93 | 1.78 | |
| ranges | 1.05 | 2.99 | 1.98 | 1.85 | 3.61 | 3.68 | 3.64 | ||
| 12 | midpoints | 0.39 | 2.10 | 1.20 | 1.09 | 3.67 | 3.17 | 3.67 | |
| ranges | 0.65 | 3.20 | 1.64 | 1.48 | 6.89 | 4.90 | 6.89 | ||
| Estimation: 2006–2011; Forecast:2012–2014 | |||||||||
| h | 2SVBMA | MMA | SAIC | SBIC | SHQ | AIC | BIC | HQ | |
| 1 | midpoints | 0.22 | 0.66 | 0.47 | 0.43 | 1.21 | 0.91 | 1.17 | |
| ranges | 0.33 | 1.03 | 0.73 | 0.66 | 2.07 | 1.61 | 1.84 | ||
| 4 | midpoints | 0.26 | 0.97 | 0.64 | 0.58 | 1.43 | 1.50 | 1.34 | |
| ranges | 0.41 | 0.97 | 0.70 | 0.64 | 1.39 | 1.27 | 1.35 | ||
| 8 | midpoints | 0.22 | 0.75 | 0.47 | 0.41 | 1.29 | 1.20 | 1.23 | |
| ranges | 0.40 | 0.80 | 0.50 | 0.47 | 1.57 | 1.22 | 1.50 | ||
| 12 | midpoints | 0.09 | 0.45 | 0.22 | 0.18 | 1.36 | 0.61 | 1.20 | |
| ranges | 0.25 | 0.42 | 0.29 | 0.28 | 0.65 | 0.56 | 0.66 | ||
| Estimation: 2007–2012; Forecast:2013–2015 | |||||||||
| h | 2SVBMA | MMA | SAIC | SBIC | SHQ | AIC | BIC | HQ | |
| 1 | midpoints | 0.07 | 0.16 | 0.14 | 0.13 | 0.38 | 0.18 | 0.34 | |
| ranges | 0.13 | 0.22 | 0.19 | 0.18 | 0.60 | 0.26 | 0.35 | ||
| 4 | midpoints | 0.09 | 0.39 | 0.24 | 0.21 | 0.58 | 0.74 | 0.58 | |
| ranges | 0.18 | 0.20 | 0.21 | 0.20 | 0.48 | 0.25 | 0.29 | ||
| 8 | midpoints | 0.17 | 0.24 | 0.24 | 0.30 | 0.35 | 0.36 | ||
| ranges | 0.37 | 0.39 | 0.46 | 0.42 | 0.44 | 1.36 | 0.79 | ||
| 12 | midpoints | 0.22 | 0.27 | 0.25 | 0.37 | 0.26 | 0.35 | ||
| ranges | 0.49 | 0.73 | 0.56 | 0.55 | 0.83 | 0.67 | 0.97 | ||
MSPE () of the recursive prediction for interval-valued crude oil prices (I).
Note: “Estimation” denotes the sample during this period used to estimate parameters, and “Forecast” denotes the sample during this period used to do out-of-sample forecasts. The best forecasts are marked by boldface, and the second best forecasts are marked by underline.
TABLE 5
| Estimation: 2008–2013; Forecast:2014–2016 | |||||||||
|---|---|---|---|---|---|---|---|---|---|
| h | 2SVBMA | MMA | SAIC | SBIC | SHQ | AIC | BIC | HQ | |
| 1 | midpoints | 0.15 | 0.30 | 0.29 | 0.27 | 0.61 | 0.34 | 0.57 | |
| ranges | 0.81 | 1.14 | 1.03 | 1.00 | 1.50 | 1.03 | 1.38 | ||
| 4 | midpoints | 0.39 | 0.47 | 0.52 | 0.50 | 0.69 | 0.65 | 0.70 | |
| ranges | 1.11 | 1.71 | 1.47 | 1.41 | 2.15 | 2.01 | 2.12 | ||
| 8 | midpoints | 1.38 | 0.93 | 0.77 | 0.87 | 2.83 | 0.43 | 3.29 | |
| ranges | 2.44 | 1.98 | 1.61 | 1.82 | 5.63 | 1.03 | 5.34 | ||
| 12 | midpoints | 0.54 | 2.45 | 1.18 | 0.88 | 1.05 | 5.84 | 5.85 | |
| ranges | 0.90 | 2.52 | 1.67 | 1.52 | 4.40 | 1.96 | 4.54 | ||
| Estimation: 2009–2014; Forecast:2015–2017 | |||||||||
| h | 2SVBMA | MMA | SAIC | SBIC | SHQ | AIC | BIC | HQ | |
| 1 | midpoints | 0.22 | 0.55 | 0.52 | 0.40 | 0.47 | 1.08 | 1.03 | |
| ranges | 1.14 | 2.37 | 2.12 | 1.71 | 1.96 | 4.45 | 4.37 | ||
| 4 | midpoints | 1.61 | 1.47 | 1.14 | 1.34 | 2.46 | 0.66 | 2.45 | |
| ranges | 2.03 | 3.46 | 2.69 | 2.48 | 5.94 | 2.75 | 6.01 | ||
| 8 | midpoints | 1.67 | 1.17 | 0.87 | 1.05 | 6.87 | 0.35 | 4.15 | |
| ranges | 1.52 | 2.92 | 2.51 | 1.91 | 2.25 | 10.06 | 9.45 | ||
| 12 | midpoints | 5.15 | 1.69 | 1.06 | 1.42 | 13.58 | 0.40 | 12.92 | |
| ranges | 0.75 | 3.93 | 2.08 | 1.81 | 8.94 | 1.51 | 8.32 | ||
MSPE () of the recursive prediction for interval-valued crude oil prices (II).
Note: “Estimation” denotes the sample during this period used to estimate parameters, and “Forecast” denotes the sample during this period used to do out-of-sample forecasts. The best forecasts are marked by boldface, and the second best forecasts are marked by underline.
Tables 4, 5 report the MSPEs of -step-ahead (1,4,8,12) forecasts for the interval-valued crude oil prices using various estimation and forecast samples. First, it is worth noticing that for the horizons of 1, 4, 8 and 12 months, the 2SVBMA method outperforms other competing methods in most cases; out of the 48 cases considered, with respect to RMSFE of midpoints and ranges, it yields the best outcomes 42 times and the second best outcomes 6 times. Intuitively, the proposed 2SVBMA method selects the important factors at the first stage and then give the optimal weights averaging across the 12 nested regression forecasts. Second, 2SVBMA based on LsoMA outperforms various model averaging and model selection methods, including MMA. One possible explanation is that leave-subject-out cross-validation is more suitable for vector autoregressive situations with heteroscedastic and auto-correlated errors. Additionally, as shown in , the approximate unbiasedness of LsoMA and its asymptotic optimality in terms of obtaining the lowest quadratic errors are established. This is why LsoMA outperforms other model averaging methods (i.e., SAIC, SBIC, and SHQ) in the second stage.
Second, the SBIC estimators always produce the second-best forecasts after the 2SVBMA estimator among all model averaging methods, while SAIC achieves higher forecast criteria than other model averaging methods. Similarly, BIC always yields best forecasts among all model selection methods, while the AIC estimator achieves higher MSFE in most cases. This happens because AIC prefers selecting the relatively complicated model, which is inappropriate for out-of-sample forecasting even though it has good in-sample fitting. A simple model may be better for out-of-sample forecasting.
Furthermore, it is shown that at the second stage, model averaging forecasts outperform model selection forecasts in almost 90% of all cases. The significant advantages of model averaging support the argument of that “model uncertainty and instability seriously impair the forecasting ability of individual predictive regression models.”
Overall, the proposed approach using interval-valued data is capable of assessing and forecasting the changes in both level and volatility. We can see from the results that forecasting with model averaging is generally better than obtaining the predictions from just one model (model selection). Since we may choose a very different model when there are small changes in the original data set, which may lead to a big change in the final conclusions, resulting in non-effective decision-making due to the unstable forecasting process. The proposed method is able to help obtain more stable decision-making when a long list of interval-valued predictors is available in a wide range of fields, for example, the daily trading strategy in the finance field.
5 Conclusion
We propose a novel 2SVBMA forecasting procedure to capture the relevant information available in the interval format and the underlying characteristics of crude oil price movements. Vector Boosting in the first stage and LsoMA in the second stage are extended to interval models with interval-valued exogenous variables. Empirical results show that our proposed approach outperforms other competing model averaging and model selection methods in terms of MSFE of midpoints and ranges.
There are some limitations and potential extensions of our study. First, more advanced optimization algorithms for interval-valued variable selection can be proposed in future work. Second, the candidate models with different structures in model averaging methods can further be developed to enhance forecasting. It would also be interesting to develop interval-based machine learning methods to improve forecast accuracy. Furthermore, the proposed methodology in this paper can be extended to the vector autoregressive (VAR) model, which can cover more applications in economics and finance.
In general, 2SVBMA provides a methodological framework for interval-valued data forecasting when there are a large number of potential predictors. For example, this methodology can be used to quantify the impact of COVID-19 pandemic on oil and gas industry. 2SVBMA can also provide implications for the post-COVID recovery management. The accurate prediction of crude oil prices will assist policy makers in understanding issues affecting different oil industry segments, and help governments be better prepared for the recovery.
6 Compliance With Ethical Standards
The authors thank a number of the participants at Symposium on Interval Data Modelling: Theory and Applications (SIDM 2019) in Beijing for their valuable comments and suggestions. This work was partially supported by National Natural Science Foundation of China (Nos. 71973116, 71988101, 72073126, 72091212), and the disciplinary funding of Central University of Finance and Economics. The authors declare no competing interests. This article does not contain any studies with human participants performed by any of the authors.
Statements
Data availability statement
The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author.
Author contributions
All three authors contributed equally to this work and the order of authorship has nothing other than alphabetical significance.
Funding
This work was partially supported by National Natural Science Foundation of China (Nos. 71973116, 71988101, 72073126, 72091212), the funding of Forecasting and Monitoring of COVID-19 in countries along "Belt and Road" and Related Economic Impacts (ANSO-SBA-2020-12), and the disciplinary funding of Central University of Finance and Economics.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
References
1
AbramsonB.FinizzaA. (1995). Probabilistic Forecasts from Probabilistic Models: a Case Study in the Oil Market. Int. J. Forecast.11, 63–72. 10.1016/0169-2070(94)02004-9
2
Álvarez-DíazM. (2019). Is it Possible to Accurately Forecast the Evolution of Brent Crude Oil Prices? an Answer Based on Parametric and Nonparametric Forecasting Methods. Empirical Econ.59, 1285–1305. 10.1007/s00181-019-01665-w
3
ArroyoJ.González-RiveraG.MatéC. (2011). “Forecasting with Interval and Histogram Data: Some Financial Applications,” in Handbook of Empirical Economics and Finance. Editors UllahA.GilesD. E. A. (New York: Chapman & Hall), 247–279.
4
BalcilarM.GuptaR.MillerS. M. (2015). Regime Switching Model of Us Crude Oil and Stock Market Prices: 1859 to 2013. Energ. Econ.49, 317–327. 10.1016/j.eneco.2015.01.026
5
BatesJ. M.GrangerC. W. J. (1969). The Combination of Forecasts. Or20, 451–468. 10.2307/3008764
6
BaurD. G.LuceyB. M. (2010). Is Gold a Hedge or a Safe haven? an Analysis of Stocks, Bonds and Gold. Financial Rev.45, 217–229. 10.1111/j.1540-6288.2010.00244.x
7
BinderK. E.PourahmadiM.MjeldeJ. W. (2018). The Role of Temporal Dependence in Factor Selection and Forecasting Oil Prices. Empirical Econ.58, 1–39. 10.1007/s00181-018-1574-9
8
BucklandS. T.BurnhamK. P.AugustinN. H. (1997). Model Selection: An Integral Part of Inference. Biometrics53, 603–618. 10.2307/2533961
9
BuhlmannP. (2006). Boosting for High-Dimensional Linear Models. Ann. Stat.34, 559–583. 10.1214/009053606000000092
10
ChaiJ.XingL.-M.ZhouX.-Y.ZhangZ. G.LiJ.-X. (2018). Forecasting the Wti Crude Oil price by a Hybrid-Refined Method. Energ. Econ.71, 114–127. 10.1016/j.eneco.2018.02.004
11
CheungY.-L.CheungY.-W.WanA. T. K. (2009). A High-Low Model of Daily Stock price Ranges. J. Forecast.28, 103–119. 10.1002/for.1087
12
De CarvalhoF. A. T.Lima NetoE. A.TenorioC. P. (2004). “A New Method to Fit a Linear Regression Model for Interval-Valued Data,” in Lecture Notes in Computer Science, K12004 Advances in Artificial Intelligence (Berlin: Springer-Verlag).
13
DingH.KimH.-G.ParkS. Y. (2016). Crude Oil and Stock Markets: Causal Relationships in Tails?Energ. Econ.59, 58–69. 10.1016/j.eneco.2016.07.013
14
DonaldS. G.ImbensG. W.NeweyW. K. (2009). Choosing Instrumental Variables in Conditional Moment Restriction Models. J. Econom.152, 28–36. 10.1016/j.jeconom.2008.10.013
15
EbrahimZ.InderwildiO. R.KingD. A. (2014). Macroeconomic Impacts of Oil price Volatility: Mitigation and Resilience. Front. Energ.8, 9–24. 10.1007/s11708-014-0303-0
16
FanJ.LiR. (2001). Variable Selection via Nonconcave Penalized Likelihood and its oracle Properties. J. Am. Stat. Assoc.96, 1348–1360. 10.1198/016214501753382273
17
FantazziniD.FomichevN. (2014). Forecasting the Real price of Oil Using Online Search Data. Ijcee4, 4–31. 10.1504/ijcee.2014.060284
18
FrankL. E.FriedmanJ. H. (1993). A Statistical View of Some Chemometrics Regression Tools. Technometrics35, 109–135. 10.1080/00401706.1993.10485033
19
García-AscanioC.MatéC. (2010). Electric Power Demand Forecasting Using Interval Time Series: A Comparison between Var and Imlp. Energy Policy38, 715–725. 10.1016/j.enpol.2009.10.007
20
González-RiveraG.LinW. (2013). Constrained Regression for Interval-Valued Data. J. Business Econ. Stat.31, 473–490. 10.1080/07350015.2013.818004
21
HamiltonJ. D. (2008). “Understanding Crude Oil Prices,”. (no. w14492).
22
HansenB. E. (2007). Least Squares Model Averaging. Econometrica75, 1175–1189. 10.1111/j.1468-0262.2007.00785.x
23
HansenB. E.RacineJ. S. (2012). Jackknife Model Averaging. J. Econom.167, 38–46. 10.1016/j.jeconom.2011.06.019
24
HeA. W. W.KwokJ. T. K.WanA. T. K. (2010). An Empirical Model of Daily Highs and Lows of West texas Intermediate Crude Oil Prices. Energ. Econ.32, 1499–1506. 10.1016/j.eneco.2010.07.012
25
HjortN. L.ClaeskensG. (2006). Focused Information Criteria and Model Averaging for the Cox hazard Regression Model. J. Am. Stat. Assoc.101, 1449–1464. 10.1198/016214506000000069
26
HjortN. L.ClaeskensG. (2003). Frequentist Model Average Estimators. J. Am. Stat. Assoc.98, 879–899. 10.1198/016214503000000828
27
HoetingJ. A.MadiganD.RafteryA. E.VolinskyC. T. (1999). Bayesian Model Averaging: a Tutorial. Stat. Sci.14, 382–417. 10.1214/ss/1009212519
28
HuZ.BaoY.ChiongR.XiongT. (2015). Mid-term Interval Load Forecasting Using Multi-Output Support Vector Regression with a Memetic Algorithm for Feature Selection. Energy84, 419–431. 10.1016/j.energy.2015.03.054
29
KangS. H.McIverR.YoonS.-M. (2017). Dynamic Spillover Effects Among Crude Oil, Precious Metal, and Agricultural Commodity Futures Markets. Energ. Econ.62, 19–32. 10.1016/j.eneco.2016.12.011
30
KilianL. (2009). Not all Oil price Shocks Are Alike: Disentangling Demand and Supply Shocks in the Crude Oil Market. Am. Econ. Rev.99, 1053–1069. 10.1257/aer.99.3.1053
31
KnightK.FuW. (2000). Asymptotics for Lasso-type Estimators. Ann. Stat.28, 1356–1378. 10.1214/aos/1015957397
32
LiD.LintonO.LuZ. (2015a). A Flexible Semiparametric Forecasting Model for Time Series. J. Econom.187, 345–357. 10.1016/j.jeconom.2015.02.025
33
LiX.MaJ.WangS.ZhangX. (2015b). How Does Google Search Affect Trader Positions and Crude Oil Prices?Econ. Model.49, 162–171. 10.1016/j.econmod.2015.04.005
34
LiaoJ.-C.TsayW.-J. (2016). Multivariate Least Squares Forecasting Averaging by Vector Autoregressive Models. Available at SSRN 2827416.
35
LiaoJ.ZongX.ZhangX.ZouG. (2019). Model Averaging Based on Leave-Subject-Out Cross-Validation for Vector Autoregressions. J. Econom.209, 35–60. 10.1016/j.jeconom.2018.10.007
36
Lima NetoE. d. A.De CarvalhoF. d. A. T. (2010). Constrained Linear Regression Models for Symbolic Interval-Valued Variables. Comput. Stat. Data Anal.54, 333–347. 10.1016/j.csda.2009.08.010
37
MaiaA. L. S.de CarvalhoF. d. A. T. (2011). Holt's Exponential Smoothing and Neural Network Models for Forecasting Interval-Valued Time Series. Int. J. Forecast.27, 740–759. 10.1016/j.ijforecast.2010.02.012
38
MaiaA. L. S.De CarvalhoF. d. A. T.LudermirT. B. (2008). Forecasting Models for Interval-Valued Time Series. Neurocomputing71, 3344–3352. 10.1016/j.neucom.2008.02.022
39
MillerJ. I.RattiR. A. (2009). Crude Oil and Stock Markets: Stability, Instability, and Bubbles. Energ. Econ.31, 559–568. 10.1016/j.eneco.2009.01.009
40
NgS.BaiJ. (2009). Selecting Instrumental Variables in a Data Rich Environment. J. Time Ser. Econom.1, 4. 10.2202/1941-1928.1014
41
PanZ.WangY.YangL. (2014). Hedging Crude Oil Using Refined Product: A Regime Switching Asymmetric Dcc Approach. Energ. Econ.46, 472–484. 10.1016/j.eneco.2014.05.014
42
QiaoK.SunY.WangS. (2019). Market Inefficiencies Associated with Pricing Oil Stocks during Shocks. Energ. Econ.81, 661–671. 10.1016/j.eneco.2019.04.016
43
RapachD. E.StraussJ. K.ZhouG. (2010). Out-of-sample Equity Premium Prediction: Combination Forecasts and Links to the Real Economy. Rev. Financ. Stud.23, 821–862. 10.1093/rfs/hhp063
44
ReboredoJ. C. (2013). Is Gold a Hedge or Safe haven against Oil price Movements?Resour. Pol.38, 130–137. 10.1016/j.resourpol.2013.02.003
45
ShinH.HouT.ParkK.ParkC.-K.ChoiS. (2013). Prediction of Movement Direction in Crude Oil Prices Based on Semi-supervised Learning. Decis. Support Syst.55, 348–358. 10.1016/j.dss.2012.11.009
46
SoučekM. (2013). Crude Oil, Equity and Gold Futures Open Interest Co-movements. Energ. Econ.40, 306–315. 10.1016/j.eneco.2013.07.010
47
SunY.HanA.HongY.WangS. (2018). Threshold Autoregressive Models for Interval-Valued Time Series Data. J. Econom.206, 414–446. 10.1016/j.jeconom.2018.06.009
48
SunY.ZhangX.HongY.WangS. (2019). Asymmetric Pass-Through of Oil Prices to Gasoline Prices with Interval Time Series Modelling. Energ. Econ.78, 165–173. 10.1016/j.eneco.2018.10.027
49
Taghizadeh-HesaryF.RasoulinezhadE.KobayashiY. (2016). Oil price Fluctuations and Oil Consuming Sectors: An Empirical Analysis of Japan. Econom. Pol. Ener. Environ. (2), 33–51. 10.3280/EFE2016-002003
50
WangX.ZhangZ.LiS. (2016). Set-valued and Interval-Valued Stationary Time Series. J. Multivariate Anal.145, 208–223. 10.1016/j.jmva.2015.12.010
51
WangY.LiuL.WuC. (2017). Forecasting the Real Prices of Crude Oil Using Forecast Combinations over Time-Varying Parameter Models. Energ. Econ.66, 337–348. 10.1016/j.eneco.2017.07.007
52
WuB.WangL.LvS.-X.ZengY.-R. (2021). Effective Crude Oil price Forecasting Using New Text-Based and Big-Data-Driven Model. Measurement168, 108468. 10.1016/j.measurement.2020.108468
53
XiongT.LiC.BaoY. (2017). Interval-valued Time Series Forecasting Using a Novel Hybrid Holti and Msvr Model. Econ. Model.60, 11–23. 10.1016/j.econmod.2016.08.019
54
XuG.WangS.HuangJ. Z. (2014). Focused Information Criterion and Model Averaging Based on Weighted Composite Quantile Regression. Scand. J. Statist41, 365–381. 10.1111/sjos.12034
55
YangW.HanA.CaiK.WangS. (2012). Acix Model with Interval Dummy Variables and its Application in Forecasting Interval-Valued Crude Oil Prices. Proced. Comp. Sci.9, 1273–1282. 10.1016/j.procs.2012.04.139
56
YangW.HanA.HongY.WangS. (2016). Analysis of Crisis Impact on Crude Oil Prices: a New Approach with Interval Time Series Modelling. Quantitative Finance16, 1917–1928. 10.1080/14697688.2016.1211795
57
YangY.GuoJ. e.SunS.LiY. (2021). Forecasting Crude Oil price with a New Hybrid Approach and Multi-Source Data. Eng. Appl. Artif. Intelligence101, 104217. 10.1016/j.engappai.2021.104217
58
YoshinoN.HesaryF. T. (2014). Monetary Policy and Oil price Fluctuations Following the Subprime Mortgage Crisis. Ijmef7, 157–174. 10.1504/ijmef.2014.066482
59
YuL.ZhaoY.TangL.YangZ. (2019). Online Big Data-Driven Oil Consumption Forecasting with Google Trends. Int. J. Forecast.35, 213–223. 10.1016/j.ijforecast.2017.11.005
60
ZaaboutiK.Ben MohamedE.BouriA. (2016). Does Oil price Affect the Value of Firms? Evidence from Tunisian Listed Firms. Front. Energ.10, 1–13. 10.1007/s11708-016-0396-8
61
ZhangX.LaiK. K.WangS.-Y. (2008). A New Approach for Crude Oil price Analysis Based on Empirical Mode Decomposition. Energ. Econ.30, 905–918. 10.1016/j.eneco.2007.02.012
62
ZhangX.LiangH. (2011). Focused Information Criterion and Model Averaging for Generalized Additive Partial Linear Models. Ann. Stat.39, 174–200. 10.1214/10-aos832
63
ZhangX.WanA. T. K.ZhouS. Z. (2012). Focused Information Criteria, Model Selection, and Model Averaging in a Tobit Model with a Nonzero Threshold. J. Business Econ. Stat.30, 132–142. 10.1198/jbes.2011.10075
64
ZhangX.YuL.WangS.LaiK. K. (2009). Estimating the Impact of Extreme Events on Crude Oil price: An Emd-Based Event Analysis Method. Energ. Econ.31, 768–778. 10.1016/j.eneco.2009.04.003
65
ZhangY.LiJ.LiuH.ZhaoG.TianY.XieK. (2020). Environmental, Social, and Economic Assessment of Energy Utilization of Crop Residue in china. Front. Energ.15, 308–319. 10.1007/s11708-020-0696-x
66
ZhaoL.ZhangX.WangS.XuS. (2016). The Effects of Oil price Shocks on Output and Inflation in china. Energ. Econ.53, 101–110. 10.1016/j.eneco.2014.11.017
67
ZhaoY.LiJ.YuL. (2017). A Deep Learning Ensemble Approach for Crude Oil price Forecasting. Energ. Econ.66, 9–16. 10.1016/j.eneco.2017.05.023
Summary
Keywords
crude oil prices forecasting, forecast combination, interval-valued time series, model averaging, vector L2-boosting
Citation
Huang B, Sun Y and Wang S (2021) A New Two-Stage Approach with Boosting and Model Averaging for Interval-Valued Crude Oil Prices Forecasting in Uncertainty Environments. Front. Energy Res. 9:707937. doi: 10.3389/fenrg.2021.707937
Received
12 May 2021
Accepted
16 July 2021
Published
19 August 2021
Volume
9 - 2021
Edited by
Farhad Taghizadeh-Hesary, Tokai University, Japan
Updates
Copyright
© 2021 Huang, Sun and Wang.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Yuying Sun, sunyuying@amss.ac.cn
This article was submitted to Sustainable Energy Systems and Policies, a section of the journal Frontiers in Energy Research
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.