ORIGINAL RESEARCH article

Front. Psychol., 06 August 2026

Sec. Quantitative Psychology and Measurement

Volume 17 - 2026 | https://doi.org/10.3389/fpsyg.2026.1787528

Guidance for data analysis in split-plot designs when parametric assumptions are violated: F-statistic, adjusted F-tests, and bootstrap-F

  • 1. Department of Psychobiology and Behavioral Sciences Methodology, University of Malaga, Málaga, Spain

  • 2. Department of Social Psychology and Quantitative Psychology, University of Barcelona, Barcelona, Spain

  • 3. Institute of Neurosciences, University of Barcelona, Barcelona, Spain

  • 4. Department of Psychology, Universidad Loyola Andalucia, Seville, Spain

  • 5. Department of Psychology, University of Oviedo, Oviedo, Spain

Abstract

Introduction:

Data analysis of split-plot designs requires normality, homogeneity of variance, and multisample sphericity. Adjusted F-tests and the bootstrap-F procedure have been proposed as alternatives to the F-statistic when these assumptions are violated. However, simulation studies have not identified the conditions under which these tests are robust.

Method:

To analyze the Type I error of the F-statistic, Greenhouse-Geisser (F-GG) and Huynh-Feldt (F-HF) adjustments, and bootstrap-F (B-F), we performed a simulation study with split-plot designs, manipulating the distribution shape, epsilon values, sample size, coefficient of sample size variation, degree of heterogeneity, and pairing of variance with group sample size.

Results:

With balanced designs, the results show that for the group effect the F-statistic and B-F are valid choices under assumption violations. For time and interaction effects, under non-sphericity, F-GG and F-HF maintain Type I error close to 5% with moderate violation of normality; both of them remain robust with 0.70 for larger violations of normality. With unbalanced designs, the behavior of all these statistics depends on the variables manipulated.

Discussion:

The paper identifies the conditions under which each procedure can be used. Overall, B-F was the most robust procedure in the majority of conditions. It can be used under moderate violations of normality, homogeneity, and sphericity with N > 20. Under more severe violations of these assumptions, larger sample size, N > 80, is needed to achieve robustness.

Introduction

The split-plot design is a mixed design that incorporates both a between-subjects or grouping factor and a within-subject factor, in which the dependent variable is registered in all participants under all conditions of the within-subject factor, but only under one level of the grouping factor. When the design involves repeated measures over several periods of time, it is conceived as longitudinal. Under the general linear model (GLM), the conventional statistical procedure is repeated measures analysis of variance (ANOVA), which uses the F-statistic to test the statistical significance associated with the null hypothesis corresponding to the different effects. This design yields two main effects, namely group and within-subject factor effects (hereinafter, time effect), as well as the interaction effect between them. For a valid statistical decision, this test requires fulfillment of the assumptions of normality, homogeneity of variance, and multisample sphericity (; ). The latter assumption implies both homogeneity of the variance-covariance matrices within each level of the between-subjects factor, and sphericity in the common variance-covariance matrix. That is, there must be sphericity for the within-subject factor for each level of the between-subjects factor.

Monte Carlo simulation studies are commonly used to examine the impact of assumption violation on the performance of statistical procedures, understood in terms of Type I error and power. Type I error is generally interpreted using liberal criterion (e.g., ; , ; ; ; ), according to which a procedure is considered robust if the Type I error rate is between 2.5% and 7.5% for a nominal alpha of 5%. The procedure is considered liberal when the Type I error rate is above 7.5%, and conservative when it is below 2.5%. The results of all the studies cited below have been interpreted according to this criterion.

Most simulation studies performed with split-plot designs have focused on the performance of procedures for time and group × time interaction effects. The group effect has typically been analyzed in studies involving one-way designs, and the empirical evidence suggests that the F-statistic for this effect is not affected by non-normality (; ; ; ; ). Regarding the homogeneity assumption, one of the most important variables is the inequality of group sizes. When the homogeneity of variances assumption is satisfied, the F-statistic is generally robust with both balanced and unbalanced designs (; ). In the case of heterogeneity, it is also robust with a balanced design (; ; ). By contrast, with unbalanced designs the robustness of the F-statistic depends on the degree of both variance heterogeneity and inequality in group sizes, as well as on the pairing of variance with group size (; ; ). The ratio of the largest to the smallest variance (variance ratio) is used as a measure of heterogeneity, while the coefficient of sample size variation (Δn) indicates the amount of inequality in group sizes. Pairing is positive when the largest group size is associated with the largest value of variance, whereas it is negative when the largest group size is associated with the smallest variance. In this regard, found that the F-statistic tended to show liberal behavior with negative pairing and a variance ratio greater than 1.5, with this behavior becoming more pronounced with increases in the variance ratio (5 and 9) and inequality of group sizes (Δn = 0.33 and 0.50). By contrast, the F-statistic tended to be conservative with positive pairing and a variance ratio greater than 2. As in the previous case, this behavior worsened with higher variance ratio and greater inequality of group sizes.

Regarding the time effect and the group × time interaction, showed that the F-statistic was robust to violations of normality, although it has a tendency to be liberal with highly skewed distributions. However, most studies have found that violation of sphericity has a serious effect on its robustness (; , ; ; Haverkamp and Beauducel, 2017, ; ; ; ). For both effects, the violation of sphericity causes the F-statistic to show liberal behavior with equal group sizes (; ; ; ), while with unequal group sizes its behavior again depends on the type of pairing; with negative pairing the F-statistic is liberal, whereas with positive pairing it is conservative (; ; , ).

A number of alternatives to the F-statistic for sphericity violations have been proposed, namely adjusted F-tests, multivariate analysis, non-parametric procedures, linear mixed models, robust statistics, and bootstrap methods (, ; , , ; ; ; ). Of these, adjusted F-tests are the most widely used as they can be implemented in most statistical packages. These procedures make the test more restrictive () by reducing the degrees of freedom of each effect by a multiplicative factor called epsilon (ε). The value of ε represents the degree of deviation from sphericity, and it ranges from its lower limit (1/K−1) to 1 (), with K being the number of repeated measures. When sphericity is satisfied, the value of ε is equal to or close to 1, whereas when it is violated, ε is further from 1 and close to its lower limit. The two most widely known and used adjusted F-tests are the Greenhouse-Geisser (F-GG; ; ; Greenhouse and Geisser, 1959) and Huynh-Feldt (F-HF;) adjustments, in which the estimators of ε are referred to as and , respectively.

There is extensive empirical evidence from simulation studies on the effect of assumption violations on the behavior of these adjusted F-tests for time and interaction effects. Regarding the time effect, and in balanced designs, both F-GG and F-HF have been found to maintain adequate control of Type I error in the presence of sphericity violation, different variance ratios between groups (e.g., 1:1:1, 1:1.5:2, and 1:3:5), and different variance-covariance structures for both normal and non-normal data (; ; ; , ; ). Both procedures have been recommended when the groups have equal sample sizes (; ).

In designs with unequal group sizes, the robustness of adjusted F-tests for the time effect depends on the degree of heterogeneity of variances, pairing of variances with group size, and the coefficient of variation of group size. If pairing is positive, adjusted F-tests are conservative, whereas if it is negative they are liberal; these behaviors become more pronounced with higher coefficients of variation of group size and smaller sample size (; , ; , ). Overall, the adjusted F-tests are not recommended when there is heterogeneity of variance.

For interaction effects the results obtained are in line with the time effects. In balanced designs, adjusted F-tests show adequate control of Type I error with normal and non-normal data under violation of homogeneity of variances and sphericity violation (; , ; ). In unbalanced designs, their behavior tends to be liberal if pairing is negative (; ; ; , ), and conservative if pairing is positive, except with a low coefficient of variation of group size and ≤ 0.75, in which case they are robust (; ; ). These results indicate that neither of the adjusted F-tests may be suitable alternative procedures under heterogeneity and unequal group size, leading some authors () to recommend balanced designs.

Another proposed alternative to ANOVA and adjusted F-tests when parametric assumptions are violated is the bootstrap method, which consists in performing a random resampling of centered data with replacement, generating a set of B bootstrap samples. The F-statistic is computed in each sample, thereby creating its empirical distribution (). The bootstrap approach was developed by as a compromise between parametric and non-parametric methods. This method does not require fulfillment of parametric assumptions and allows researchers to obtain more accurate estimators in situations of assumption violations (; ; ). developed bootstrap-F (B-F) and assessed its robustness through a Monte Carlo simulation study in a one-way design with four repeated measures. They manipulated different sample sizes (from 10 to 60), ε values (from 0.48 to 1), and different values of skewness (γ1) and kurtosis (γ2) so as to generate non-normal distributions corresponding to slight (γ1 = 1, γ2 = 0.75), moderate (γ1 = 1.75, γ2 = 3.00), and severe deviations from normality (γ1 = 3.00, γ2 = 21.00). The results showed that under simultaneous violations of normality and sphericity, B-F is robust, even with small sample sizes, although it may become conservative in some scenarios with ε = 1. Similar results have recently been reported by for one-way designs under a wider variety of scenarios of non-normality and non-sphericity, although it was also suggested that the sample size should be larger than 20–25 with severe violations of parametric assumptions.

described the bootstrap procedure for split-plot designs. These authors studied the behavior of B-F in both balanced and unbalanced 3 × 4 split-plot designs, manipulating the total sample size (N = 30, 45, and 60), coefficients of sample size variation (Δn = 1 and 0.33), degree of heterogeneity of the variance-covariance matrix (1:1:1, 1:3:5), type of pairing (null, positive, and negative), ε values (ε = 0.50, 0.70, and 1), and distribution shape (same distributions as ). Regarding the group effect, the results showed that B-F was generally robust, except in two conditions of extreme violation in which it showed conservative behavior: (a) N = 30, positive pairing, severe deviation from normality (γ1 = 3.00, γ2 = 21.00), and severe deviation from sphericity (ε = 0.50); and (b) N = 30, positive pairing, severe non-normality, while the covariance matrices differed across groups and were not proportional, with different levels of sphericity across groups (ε = 1, 0.75, and 0.50). For the time effect, the results showed that B-F adequately controlled Type 1 error with normal data. However, B-F was conservative with severe deviation from normality and the sphericity assumption, regardless of the type of pairing. Regarding the interaction effect, B-F was again robust with normal distributions. With non-normal distributions, B-F was conservative with severe deviation from normality and ε = 0.70 and 1. In a subsequent study, reported adequate control of Type I error for B-F with a 3 × 4 split-plot design (balanced and unbalanced), N = 30 and 45, an ε value of 0.50, equal and unequal variance-covariance matrices (1:3:5 and 1:5:9), and normal and non-normal distributions (double exponential, exponential, and lognormal). Overall, the empirical evidence indicates that B-F is robust to a wide variety of situations in which the assumptions of split-plot designs are violated, and hence it can be considered a suitable alternative for analysis in these situations.

In general, studies show that the F-statistic is sensitive to homogeneity and sphericity assumptions and that F-adjusted tests may be a good alternative for time and interaction effects in balanced designs. However, in unbalanced designs, adjusted F-tests do not correct for the effect of non-sphericity and they can be conservative or liberal, depending on the conditions of heterogeneity of variances and group sample sizes. The conclusion, therefore, is that they are not suitable alternatives in unbalanced designs. However, simulation studies have not identified the conditions under which these tests could be used with guarantees of robustness, nor do they clarify how the non-normality of data influences Type I error under the different scenarios of non-sphericity and heterogeneity. This contrasts with actual practice in research, insofar as adjusted F-tests are systematically used in split-plot designs without considering the results of simulation studies. As for B-F, this procedure seems to show adequate performance and it can be considered a promising alternative. However, Monte Carlo studies are very limited, and its behavior needs to be analyzed in a larger number of scenarios.

The aim of the present study is to provide new evidence on the behavior of the F-statistic, adjusted F-tests, and B-F with split-plot designs in the presence of assumption violations, clarifying the conditions in which each of these procedures may be suitable. Simulation studies with split-plot designs have frequently used 3 × 4 designs (e.g., ; ; ; ; , ). However, research practice reveals that the most frequently used number of levels is 2 for the grouping factor and 2 or 3 for the within-subject factor (; ). Therefore, we conduct an exhaustive Monte Carlo simulation study with a 2x3 split-plot design and 5,250 conditions, manipulating the following variables: Distribution shape (normal, as well as slight, moderate, severe, and extreme deviation from normality), epsilon values (from = 0.50 to 1), total sample size (N from 20 to 200), coefficient of group size variation (Δn = 0, 0.16, 0.33, and 0.50), ratio of the variance-covariance matrices between the two groups (1:1, 1:1.5, 1:2, and 1:5), and type of pairing of variance with group size (null, positive, and negative). Our goal is to identify the conditions under which the F-statistic, F-GG, F-HF, and B-F can be used in 2 × 3 designs, providing guidelines for applied researchers regarding the use of these procedures.

Materials and methods

A Monte Carlo simulation study was conducted using the PROC GLM module of SAS 9.4, together with the interactive matrix language (IML) module. A series of ad hoc macros were created for data generation. Normal data were generated using the Cholesky transformation of the variance-covariance matrix, while non-normal data were generated using procedure. To simulate data with different degrees of sphericity violation, a series of unstructured variance-covariance matrices were generated with different values of .

The variables manipulated were as follows:

  • Shape of the distribution: The normal distribution and four non-normal distributions were used, with different values of skewness and kurtosis coefficients. Non-normal distributions were grouped into slight (γ1 = 0.4, γ2 = 0.8), moderate (γ1 = 1, γ2 = 1.5), severe (γ1 = 1.41, γ2 = 3), and extreme (γ1 = 2.31, γ2 = 8) deviation from normality. found that 80% of real data showed skewness and kurtosis values ranging between −1.25 and 1.25. Distributions with slight and moderate deviations from normality were selected based on this finding. Distributions with severe and extreme deviations were included to represent larger departures, which have typically been used in simulation studies involving repeated measures (, , ).

  • Epsilon values (): The values of were 0.50, 0.60, 0.70, 0.80, 0.90, and 1.

  • Total sample size: Seven sample sizes were considered, including small, medium, and large samples: 20, 40, 60, 80, 100, 150, and 200.

  • Coefficient of sample size variation (Δn). This represents the amount of inequality in group sizes and was calculated by dividing the standard deviation of the group size by its mean. The values of Δn were 0 (balanced design), 0.16 (low value), 0.33 (medium value), and 0.50 (high value).

  • Ratio of the variance-covariance matrices between the two groups (variance ratio, VR): (a) homogeneity of variance, 1:1; (b) heterogeneity of variance, 1:1.5, 1:2, and 1:5. In this scenario, the variance-covariance matrix of the second group could be, respectively, 1.5, 2, or 5 times larger than that of the first group.

  • Pairing of variance with group sample size. The pairings used were: null when there are equal group sizes; positive when the largest group size is associated with the largest value of variance; and negative when the largest group size is associated with the smallest variance.

Five thousand replications were performed for each of the conditions. We recorded the Type I error rate of the F-statistic, F-GG, F-HF, and B-F, which represents the proportion of rejection of the null hypothesis when it is true. In line with other researchers and to facilitate comparison between studies, liberal criterion was used to evaluate robustness, and hence a procedure was considered robust if the Type I error rate was between 2.5% and 7.5% for a nominal alpha of 5%. A procedure was considered liberal when the Type I error rate was above 7.5%, and conservative when it was below 2.5%.

Results

Balanced designs

Group effect

Figure 1 shows the results for balanced designs and the group effect (840 conditions). Type I error rates for the F-statistic and B-F are always within the bounds [2.5%, 7.5%], indicated by dashed lines. Specifically, Type I errors for these procedures are between 3.58% and 7.32% (M = 5.07, SD = 0.47) and between 2.60% and 6.36% (M = 4.83, SD = 0.56), respectively. Therefore, both the F-statistic and B-F control the Type I error rate under violations of normality and heterogeneity of variances in the conditions manipulated here.

Figure 1

Time effect

Results for the time effect are shown in Figure 2. With = 0.50 and 0.60, the F-statistic is liberal, with Type I error rates as high as 13.92%. With = 0.70, the F-statistic is close to 7.5%, and with ≥ 0.80 it controls Type I error. With ≤ 0.60, F-GG and F-HF tend to be liberal with severe and extreme deviation from normality and small sample size. With 0.70, both statistics are robust and close to 5%. F-GG and F-HF get closer to 5% than does the F-statistic as departs from 1. B-F is robust in practically all cases, except with = 1, a variance ratio of 5, and N = 20, in which case it becomes conservative. Type I error rates for B-F range from 2.38% to 7.26% (M = 5.01, SD = 0.56).

Figure 2

Group × time interaction

Figure 3 shows the results for the group × time interaction. The interaction effect of the F-statistic is affected mainly by the violation of sphericity, with a maximum Type I error of 11.80%. It exhibits a tendency toward liberality with ≤ 0.70. However, F-GG and F-HF are robust in all cases, except with = 0.50, extreme deviation from normality, a variance ratio of 5, and N = 20. Type I error rates for these two procedures range from 3.24% to 7.86% (M = 4.96, SD = 0.45), and from 3.52% to 7.90% (M = 5.05, SD = 0.43), respectively. B-F is also robust, except with = 1, extreme violation of normality, and N = 20, in which case it becomes conservative. Type I error rates for B-F range from 2.10% to 6.70% (M = 4.81, SD = 0.58).

Figure 3

Unbalanced designs with homogeneity of variance

Group effect

Results for the conditions in which homogeneity of variance is fulfilled (VR = 1; null pairing between variance and sample sizes) are shown in Figure 4 (630 conditions). In this scenario, the F-statistic for the group effect is robust in all conditions of non-normality and non-sphericity, with Type I error from 3.86% to 5.92% (M = 4.92, SD = 0.33). B-F is also generally robust, except for some conditions with N = 20 when the group sizes are very unequal (Δn = 0.50), in which case it shows a tendency to be liberal and is outperformed by the F-statistic. Type I error rates for B-F range from 3.10% to 7.54% (M = 5.35, SD = 0.63).

Figure 4

Time effect

Figure 5 shows results for the time effect. The F-statistic is liberal for ≤ 0.60 and robust with ≥ 0.70. F-GG and F-HF are generally robust, except with extreme deviation from normality and sphericity: = 0.50 and N ≤ 60, and = 0.60 and N =20, in which case they are liberal. B-F is also generally robust, except with ≤ 0.60, very unequal group sizes (Δn = 0.50), and small sample size (N ≤ 40), in which case it tends to be liberal. The greater the deviation from normality, the larger the sample size required to achieve robustness (e.g., N > 20 for moderate and severe and N > 40 for extreme deviation from normality).

Figure 5

Group × time interaction

Figure 6 displays the results for the group × time interaction, which are in line with those for the time effect. The F-statistic for the interaction effect is liberal for ≤ 0.60 and robust with ≥ 0.70, although it is not until = 0.80 that this statistic is shown to be close to 5%. F-GG and F-HF are robust in all conditions manipulated. Type I error rates for B-F are always within the bounds [2.5%, 7.5%], the maximum value reached being 7.48%.

Figure 6

Unbalanced designs with heterogeneity: positive pairing

Group effect

Table 1 and Figure 7 display the results for the group effect when there is no homogeneity of variance (VR > 1) and the largest group size is associated with the largest value of variance (1890 conditions). We omitted non-sphericity because it is not a relevant variable for group effect, as already seen for the homogeneity condition with unbalanced designs. In this scenario, Type I error shows approximately the same pattern across distribution shapes and sample sizes, with the variables that most affect it being the VR and the inequality of group sizes (Δn). With VR = 1.5, the F-statistic is robust in all conditions of Δn. With VR = 2 and 5, it is robust with Δn = 0.16, but is conservative with greater inequality of group sizes. For example, with Δn = 0.50 it may reach a minimum Type I error of 0.40%. B-F remains within the bounds [2.5%, 7.5%] to be considered robust, reaching a minimum and maximum Type I error of 2.64% and 6.90%, respectively.

Table 1

VRΔnFB-F
PairingMinMaxMSDMinMaxMSD
Positive1.50.163.685.204.290.292.845.944.800.57
0.332.964.123.640.272.645.684.960.43
0.502.463.883.010.234.206.905.490.40
20.163.444.543.930.263.005.884.770.52
0.332.243.822.950.293.065.524.830.42
0.501.583.522.160.263.466.365.220.44
50.162.685.403.290.484.145.564.820.32
0.330.984.461.810.563.945.224.670.29
0.500.403.060.870.393.645.484.720.36
Total0.162.685.403.840.552.845.944.800.48
0.330.984.462.800.852.645.684.820.40
0.500.403.882.010.933.466.905.150.51

Descriptive statistics for Type I error (in percentages) for the group effect in unbalanced designs and positive pairing (heterogeneity of variance: variance ratio greater than 1) for the F-statistic and B-F as a function of variance ratio (VR), and coefficient of sample size variation (Δn).

Type I error rates < 2.5% are in italics (conservative). n = 210 conditions for each cell. N = 630 conditions for total cells.

Figure 7

Time effect

Figures 810 show results for the time effect. The behavior of the F-statistic depends on all the variables manipulated, becoming liberal in some conditions with a low value of , especially with extreme violation of normality, and conservative with high values of . Overall, it is robust for all distributions and VRs for medium values of and Δn = 0.16. By contrast, F-GG and F-HF show conservative behavior as non-normality, the variance ratio, inequality of group sizes, and the value increase. With VR = 5, these procedures are only robust with a low value of and Δn = 0.16, except with extreme deviation and N = 20. Regarding B-F, it generally shows robust behavior for the time effect with positive pairing, although it may become liberal in some conditions with Δn = 0.50, and conservative with a high value of . As a general rule, it is robust in all conditions with N > 20.

Figure 8

Figure 9

Figure 10

Group × time interaction

Figures 1113 show the results for the group × time interaction. The F-statistic shows liberal behavior in some conditions with the lowest value of , and a tendency to be conservative with higher values of , becoming more conservative with higher VR, greater deviation from normality, and greater group inequality. As for the time effect, it is generally robust for all distributions and VRs for medium values of and Δn = 0.16. F-GG and F-HF are robust under some scenarios with Δn = 0.16–0.33, and tend to be conservative as non-normality, the variance ratio, inequality of group sizes, and the value increase. Overall, with VR = 5, F-GG and F-HF are generally conservative under all distributions, Δn, and sample size conditions. In contrast to these statistics, B-F is generally robust with positive pairing for the interaction effect. It is only conservative in a few cases, with ≥ 0.90, N = 20, and extreme violation of normality. As a general rule, it is robust in all conditions with N > 20.

Figure 11

Figure 12

Figure 13

Unbalanced designs with heterogeneity: negative pairing

Group effect

Table 2 and Figure 14 show the results when there is no homogeneity of variance (VR > 1) and the largest group size is associated with the smallest value of variance (1890 conditions). Type I error of the F-statistic is more affected under a high variance ratio and large inequality of group sizes. As can be observed in Table 2, with VR = 1.5 and 2, it is only robust with Δn = 0.16. With VR = 5, it is liberal in all conditions of Δn, reaching a Type I error as large as 19.70%. Conversely, B-F is robust with VR = 1.5 and 2, with Δn = 0.16 and 0.33, and becomes liberal with Δn = 0.50 and small sample sizes (N = 20). With VR = 5, it tends to be liberal as the deviation from normality and inequality of group sizes increase, especially with small sample size. In this case, B-F reaches robustness with N > 20 for Δn = 0.16, N > 60 for Δn = 0.33, and N > 80 for Δn = 0.50. Its maximum Type I error is 13.78%.

Figure 14

Table 2

VRΔnFB-F
PairingMinMaxMSDMinMaxMSD
Negative1.50.165.106.385.740.283.625.764.990.33
0.335.227.926.510.454.506.285.350.34
0.506.448.307.590.344.888.946.330.94
20.165.786.986.370.264.085.805.100.28
0.336.729.147.840.454.546.905.510.43
0.508.7410.649.820.394.9610.146.571.19
50.167.3411.808.380.773.988.865.370.69
0.3311.0814.3812.160.564.5610.645.800.96
0.5016.0619.7017.350.724.8013.786.861.79
Total0.165.1011.806.831.233.628.865.150.49
0.335.2214.388.842.464.5010.645.550.66
0.506.4419.7011.594.214.8013.786.591.37

Descriptive statistics for Type I error (in percentages) for the group effect in unbalanced designs and negative pairing (heterogeneity of variance: variance ratio greater than 1) for the F-statistic and B-F as a function of variance ratio (VR), and coefficient of sample size variation (Δn).

Type I error rates < 2.5% are in italics (conservative) and >7.5 are in bold (liberal). n = 210 conditions for each cell. N = 630 conditions for total cells.

Time effect

Figures 1517 show the results for the time effect. The F-statistic is only robust with VR = 1.5 and with Δn = 0.16 and ≥ 0.90 for all distributions, except for extreme deviation, when it is only robust with = 1. It is generally liberal for the remaining values of . Its behavior is more liberal the higher the VR, the higher the Δn, and the greater the violation of normality. Its Type I error rate can reach 27.34%. The same behavior, but less pronounced, is shown by F-GG and F-HF, which can reach Type I error rates of 23.24% and 23.50%, respectively. These procedures are only robust in some scenarios with VR ≤ 2 and Δn = 0.16. B-F also shows a tendency toward liberality in some conditions with low values of and small sample size (maximum Type I error of 13.22%). Its behavior worsens with high values of variance ratio, more inequality of group sizes, and greater deviation from normality. As these scenarios worsen, larger sample sizes are required to achieve robustness. There is also a tendency to be conservative with small sample size and = 1, especially with extreme deviation from normality.

Figure 15

Figure 16

Figure 17

Group × time interaction

Figures 1820 show the results for the group × time interaction. The F-statistic for this interaction is liberal under almost all conditions, especially with more unequal group sizes and higher VRs. Its Type I error can reach 26.54%. It is only robust with VR = 1.5 and with Δn = 0.16 and ≥ 0.90 for all distributions. The same tendency toward liberality is shown by F-GG and F-HF, albeit less pronounced. These procedures can yield Type I error rates of 23.34% and 23.58%, respectively, and they are only robust in some scenarios with VR ≤ 2 and Δn = 0.16. Finally, B-F is robust with VR ≤ 2 and Δn = 0.16–0.33 with N > 20, showing a tendency toward liberality in some conditions with low values of , small sample size, and greater inequality of group size. The lower the value of , the higher the VR, the greater the inequality of group sizes, and the greater the deviation from normality, the larger the sample size that is needed to achieve robustness. For example, with = 0.50, VR = 5, and Δn = 0.50 under extreme deviation from normality, B-F needs N > 60 to reach robustness. Its maximum Type I error is 14.50%.

Figure 18

Figure 19

Figure 20

Discussion

The aim of this study was to provide new evidence on the behavior of the F-statistic, adjusted F-tests, and the bootstrap method (B-F) in split-plot designs in the presence of parametric assumption violations, clarifying the conditions in which each of them may be a suitable procedure. To this end, we conducted a Monte Carlo simulation study and manipulated the following variables: Distribution shape (normal, as well as slight, moderate, severe, and extreme deviation from normality), sphericity by epsilon values (from = 0.50 to 1), total sample size (N from 20 to 200), coefficient of group size variation (Δn = 0, 0.16, 0.33, and 0.50), ratio of the variance-covariance matrices between the two groups (1:1, 1:1.5, 1:2, and 1:5), and type of pairing of variance with group sample size (null, positive, and negative). The results are discussed below according to the type of design and the variables manipulated.

Balanced designs

In balanced designs, and for the group effect, both the F-statistic and B-F control Type I error under heterogeneity of variance and non-normality, in the conditions studied here. Logically, Type I error is not affected by violation of sphericity, as this is not a relevant assumption for the between-subjects factor. These results are in line with previous research showing the robustness of the F-statistic in balanced designs (; ; ).

For time and interaction effects, the F-statistic is liberal with low epsilon values ( ≤ 0.70), reflecting the effect of sphericity violation (, ; ; ; ; ). The F-statistic performs worse than F-GG and F-HF, which are robust, except with severe and extreme normality violation, ≤ 0.60, and small samples, in which case they are liberal. Although both adjusted F-tests have been recommended when the groups have equal sample sizes (; ), our results show that, in general, they may be used for ≥ 0.60 and N > 20. By contrast, B-F is robust in almost all conditions, although it tends to be conservative with extreme violation of normality, = 1, and N = 20. This tendency to be conservative with high values of has been reported previously (; ; ).

Unbalanced designs with homogeneity of variance-covariance matrices

For the group effect, and in unbalanced designs (Δn > 0) when the homogeneity assumption is satisfied, the F-statistic is robust and is unaffected by non-normality and non-sphericity. Other studies have concluded the same (; ), which is logical since the homogeneity assumption is satisfied. In these cases, B-F is also generally robust, although it requires N > 20 if there are very unequal group sizes (Δn = 0.50).

For time and interaction effects, the F-statistic is liberal for ≤ 0.60 and robust with ≥ 0.70, although its Type I error is closer to 5% when ≥ 0.80. F-GG and F-HF are generally robust, except with extreme deviation from both normality and sphericity that they are liberal. In these cases, for both time and interaction effects, we need N > 60 if = 0.50, and N > 20 if = 0.60. B-F shows a similar behavior, and overall it requires N > 40 if ≤ 0.60.

Unbalanced designs with heterogeneity of variance-covariance matrices and positive pairing

When the homogeneity assumption is not satisfied, the behavior of the F-statistic for the group effect depends on the pairing between variances and group sizes. With positive pairing, the F-statistic tends to be conservative if the group sizes are very unequal, with the situation getting worse the greater the heterogeneity of variance. By contrast, B-F remains within the bounds [2.5%, 7.5%] to be considered robust.

For time and interaction effects, the F-statistic tends to be liberal in some conditions with the lowest value of , although there is a tendency for it to be conservative with higher values of . In general, it is robust to all distributions and VRs for medium values of and Δn = 0.16. F-GG and F-HF tend to be conservative, a tendency that is aggravated with greater heterogeneity of variance, more unequal group sizes, and higher values. These results are in line with previous research (; , ; , ). According to our results, F-GG and F-HF can, as a general rule, be used in all distributions when: (a) VR = 1.5, Δn = 0.16 and 0.33, and N > 20; (b) VR = 2, 0.60 ≤ ≤ 0.90, and Δn = 0.16; and (c) VR = 5, ≤ 0.60, Δn = 0.16, and N > 20. As for B-F, it generally shows robust behavior, except in several scenarios with extreme deviation from normality, very unequal group sizes (Δn = 0.50), and small sample size (N = 20), in which cases it may be liberal with a low value or conservative when the value of is high. Overall, B-F can be used with N > 20. These results extend knowledge about the behavior of B-F, highlighting its worsening performance with very unequal group sizes. It should be noted that previous research only included a maximum value for the group size coefficient of variation of 0.33 (, ).

Unbalanced designs with heterogeneity of variance-covariance matrices and negative pairing

With negative pairing, both the F-statistic and B-F tend to be liberal for the group effect, reaching Type I error rates of 19.70% and 13.78%, respectively. If group sizes are not very unequal (Δn = 0.16), the F-statistic is robust with VR ≤ 2; otherwise, it tends to be liberal, regardless of distribution shape, group size or variance ratio. Regarding B-F, the results also indicate a tendency toward liberality, although its behavior is better than that of the F-statistic. B-F is robust with VR ≤ 2 and Δn ≤ 0.33, and with VR ≤ 2, Δn = 0.50, and N > 20. With VR = 5, a large sample size is needed to achieve robustness, depending on the inequality of group sizes. In the worst scenario, N > 80 is required to reach robustness.

Regarding time and interaction effects, the F-statistic and the two adjusted-F tests are generally liberal, with all three reaching Type I error rates above 20%. The F-statistic is only an analytic option with VR = 1.5, = 1, and Δn = 0.50. The two adjusted-F tests are only robust in some conditions with VR ≤ 2 and Δn = 0.16. When there is large inequality in group sizes, none of these three statistics should be used. The tendency to liberality has been previously reported, also highlighting the effect of inequality of group sizes (; , ; ; , ). B-F outperforms all these statistics, but consistently shows a tendency toward liberality in some conditions with low values of , small sample size, and large inequality of group sizes. In these cases, a larger sample size is needed to achieve robustness. As for the group effect, with extreme conditions B-F needs N > 80 to reach robustness.

Conclusions: practical recommendations

The results of the present study serve to identify the conditions under which the F-statistic, adjusted F-tests, and B-F can be applied with a guarantee of robustness, as well as those conditions under which robustness is threatened. A guideline to the analysis of 3 × 2 split-plot designs that may be useful to applied researchers is provided below.

Table 3 summarizes the conditions under which the analyzed statistics can be used in balanced designs. For the group effect, the F-statistic and B-F are valid choices for analysis, regardless of whether the assumptions of normality and homogeneity are satisfied independently or in combination. For time and interaction effects, the F-statistic, as expected, is also suitable if sphericity is satisfied, whereas B-F would require N > 20 if = 1. In the event of sphericity violation, F-GG and F-HF maintain Type I error close to 5% with moderate violation of normality. For distributions with larger deviation from normality, both adjusted F-tests can be used with ≥ 0.70, but with lower values a large sample size is required ( = 0.50, N > 60; = 0.60, and N > 20). As a rule of thumb, B-F can be used to test all three effects in all conditions studied here, provided that N > 20.

Table 3

Balanced design, pairing null, all Δn
VREffectNormal–Moderate (γ1 ≤ 1, γ2 ≤ 1.5)Severe (γ1 = 1.41, γ2 = 3)Extreme (γ1 = 2.31, γ2 = 8)General rule
1, 1.5, 2, 5GFF
B-FB-F
T, G × TF, F, F,
F-GG, F-HFF-GG, F-HF, , N>20F-GG, F-HF, , N>60F-GG, F-HF, , N>60
F-GG, F-HF, F-GG, F-HF, , N>20F-GG, F-HF, , N>20
F-GG, F-HF, F-GG, F-HF,
B-FB-FB-F, B-F,
B-F, , N>20B-F, , N>20

Conditions in which the F-statistic, adjusted F-tests, and B-F can be used with balanced designs.

G, group effect; T, time effect; G × T, group × time effect; Δn, coefficient of sample size variation; VR, variance ratio; γ1, skewness coefficient; γ2, kurtosis coefficient; N, total sample size; , Greenhouse–Geisser epsilon; F, F-statistic; F-GG, Greenhouse–Geisser adjusted F-test; F-HF, Huynh–Feldt adjusted F-test; B-F, bootstrap-F.

Table 4 summarizes the conditions in which the analyzed statistics can be used in unbalanced designs with homogeneity of variance. The F-statistic can be used for the group effect, regardless of non-normality and unequal group sizes, as well as for the time and interaction effect when ≥ 0.70. B-F may also be a suitable alternative for all three effects, but for the group effect it requires N > 20 if there is large inequality of group sizes, while for the time and interaction effects, N > 40 is needed if ≤ 0.60. As for balanced designs, and in the event of sphericity violation, F-GG and F-HF maintain Type I error close to 5% with moderate violation of normality. For distributions with larger deviation from normality, both can be used with ≥ 0.70, but with lower values a large sample size is required ( = 0.50, N > 60; = 0.60, N > 20).

Table 4

Unbalanced design, homogeneity of variance-covariance matrices
VREffectNormal–Severe (γ1 ≤ 1.41, γ2 ≤ 3)Extreme (γ1 = 2.31, γ2 = 8)General rule
1GFF
B-F, Δn = 0.16–0.33B-F, Δn = 0.16–0.33
B-F, Δn = 0.50, N>20B-F, Δn = 0.50, N>20
T, G × TF, F,
F-GG, F-HFF-GG, F-HF, , N>60F-GG, F-HF, , N>60
F-GG, F-HF, , N>20F-GG, F-HF, , N>20
F-GG, F-HF, F-GG, F-HF,
B-F, Δn = 0.16–0.33B-F, , N>40B-F, , N>40
B-F, Δn = 0.50, N>20B-F, B-F,

Conditions in which the F-statistic, adjusted F-tests, and B-F can be used for unbalanced designs with homogeneity of variance.

G, group effect; T, time effect; G × T, group × time effect; Δn, coefficient of sample size variation; VR, variance ratio; γ1, skewness coefficient; γ2, kurtosis coefficient; N, total sample size; , Greenhouse–Geisser epsilon; F, F-statistic; F-GG, Greenhouse–Geisser adjusted F-test; F-HF, Huynh–Feldt adjusted F-test; B-F, bootstrap-F.

Table 5 summarizes the conditions in which the analyzed statistics can be used in unbalanced designs with heterogeneity of variance and positive pairing. For the group effect the F-statistic can generally only be used when the groups are of similar size (Δn = 0.16), while for the time and interaction effects, it may, as a rule of thumb, be used with Δn = 0.16 and 0.70 ≤ ≤ 0.80. F-GG and F-HF can, as a general rule, be used for the time and interaction effects in all distributions when: (a) VR = 1.5, Δn = 0.16 and 0.33, and N > 20; (b) VR = 2, 0.60 ≤ ≤ 0.90, and Δn = 0.16; and (c) VR = 5, ≤ 0.60, Δn = 0.16, and N > 20. By contrast, B-F is the procedure that best controls the Type I error rate, and it can be used for all three effects with N > 20. Therefore, it is considered the best choice for analysis in cases where the largest group size has the largest variance value.

Table 5

Unbalanced design, positive pairing
VREffectNormal–Severe (γ1 ≤ 1.41, γ2 ≤ 3)Extreme (γ1 = 2.31, γ2 = 8)General rule
1.5+GFF
B-FB-F
T, G × TF, F,
F-GG, F-HF, Δn = 0.16–0.33F-GG, F-HF, , N > 20F-GG, F-HF, Δn = 0.16–0.33, N > 20
F-GG, F-HF, , Δn = 0.16–0.33
F-GG, F-HF, , Δn = 0.16–0.33, N > 20
B-FB-F, N > 20B-F, N > 20
2+GF, Δn = 0.16F, Δn = 0.16
B-FB-F
T, G × TF, , Δn = 0.16–0.33F, , Δn = 0.16–0.33F, , Δn = 0.16–0.33
F, , Δn = 0.16F, , Δn = 0.16F, , Δn = 0.16
F-GG, F-HF, Δn = 0.16F-GG, F-HF, , Δn = 0.16F-GG, F-HF, , Δn = 0.16
B-FB-F, N > 20B-F, N > 20
5+GF, Δn = 0.16F, Δn = 0.16
B-FB-F
T, G × TF, , Δn = 0.16F, , Δn = 0.16F, , Δn = 0.16
F-GG, F-HF, , Δn = 0.16F-GG, F-HF, , Δn = 0.16, N > 20F-GG, F-HF, , Δn = 0.16, N > 20
B-FB-F, B-F, N > 20
B-F, , N > 20

Conditions in which the F-statistic, adjusted F-tests, and B-F can be used for unbalanced designs with heterogeneity of variance and positive pairing.

G, group effect; T, time effect; G × T, group × time effect; Δn, coefficient of sample size variation; VR, variance ratio; γ1, skewness coefficient; γ2, kurtosis coefficient; N, total sample size; , Greenhouse–Geisser epsilon; F, F-statistic; F-GG, Greenhouse–Geisser adjusted F-test; F-HF, Huynh–Feldt adjusted F-test; B-F, bootstrap-F.

In the majority of cases with positive pairing, the F-statistic and adjusted F-tests tend to be conservative. When a test is conservative, there is a greater probability of accepting the null hypothesis when it is false. Therefore, researchers can be sure in these scenarios that the statistical decision is correct if the null hypothesis is rejected.

Table 6 summarizes the conditions in which the analyzed statistics can be used in unbalanced designs with heterogeneity of variance and negative pairing. The F-statistic can only be used for the three effects if the group sizes are very similar (Δn = 0.16), the variance ratio is close to 1.5, and sphericity is fulfilled ( = 1). Similarly, the two adjusted F-tests may only be used with similar group sizes (Δn = 0.16) and: (a) a variance ratio of 1.5 and ≥ 0.70; or (b) a variance ratio of 2, 70 ≤ ≤ 0.80, and N > 20. These two statistics, as well as the F-statistic, have serious limitations in conditions of negative pairing with very unequal variances, where they tend to be liberal. In these cases, researchers will more likely reject the null hypothesis when it is in fact true. Therefore, one can only be sure that the statistical decision is correct if the results lead to acceptance of the null hypothesis. Regarding B-F, it can be used for all three effects provided the sample size is adjusted based on the degree of inequality of the group sizes and the severity of the violation of both homogeneity and normality. As a rule of thumb, it can be used with: (a) Δn = 0.16, N > 20; (b) Δn = 0.33, N > 60; and (c) Δn = 0.50, N > 80. Therefore, if a sample size greater than 80 is available, B-F can be used with guarantees of robustness, regardless of assumption violations.

Table 6

Unbalanced design, negative pairing
VREffectNormal (γ1 = 0, γ2= 0)Slight (γ1 = 0.4, γ2= 0.8)Moderate (γ1 = 1, γ2= 1.5)Severe (γ1 = 1.41, γ2= 3)Extreme (γ1 = 2.31, γ2= 8)General rule
1.5–GF, Δn = 0.16F, Δn = 0.16
B-F, Δn = 0.16–0.33B-F, Δn = 0.16–0.33
B-F, Δn = 0.50, N > 20B-F, Δn = 0.50, N > 20
T, GxTF, ≥ 0.90 Δn = 0.16F, = 1, Δn = 0.16F, = 1, Δn = 0.16
F-GG, F-HF, Δn = 0.16F-GG, F-HF, ≥ 0.60, Δn = 0.16F-GG, F-HF, ≥ 0.70, Δn = 0.16F-GG, F-HF, ≥ 0.70, Δn = 0.16
B-F, Δn = 0.16–0.33B-F, Δn = 0.16–0.33
B-F, Δn = 0.50, N > 20B-F, Δn = 0.50, N > 20
2–GF, Δn = 0.16F, Δn = 0.16
B-F, Δn = 0.16–0.33B-F, Δn = 0.16–0.33B-F, Δn = 0.16–0.33
B-F, Δn = 0.50, N > 20B-F, Δn = 0.50, N > 80B-F, Δn = 0.50, N > 80
T, GxTF-GG, F-HF, Δn = 0.16, N > 20F-GG, F-HF, ≥ 0.70, Δn = 0.16F-GG, F-HF, 0.70 ≤ ≤ 0.80, Δn = 0.16F-GG, F-HF, 0.70 ≤ ≤ 0.80, Δn = 0.16, N > 20
B-F, Δn = 0.16–0.33B-F, ≤ 0.60, Δn = 0.16–0.33, N > 20B-F, Δn = 0.16–0.33, N > 20
B-F, Δn = 0.50, N > 20B-F, ≤ 0.60, Δn = 0.50, N > 80B-F, Δn = 0.50, N > 80
B-F, 0.70 ≤ ≤ 0.90
B-F, = 1, N > 20
5–GB-F, Δn = 0.16–0.33B-F, Δn = 0.16–0.33B-F, Δn = 0.16B-F, Δn = 0.16, N > 20B-F, Δn = 0.16, N > 20
B-F, Δn = 0.50, N > 20B-F, Δn = 0.50, N > 40B-F, Δn = 0.33, N > 20B-F, Δn = 0.33, N > 60B-F, Δn = 0.33, N > 60
B-F, Δn = 0.50, N > 60B-F, Δn = 0.50, N > 80B-F, Δn = 0.50, N > 80
T, GxTB-F, Δn = 0.16–0.33B-F, Δn = 0.16–0.33B-F, Δn = 0.16B-F, ≤ 0.60, Δn = 0.16–0.33, N > 60B-F, N > 60
B-F, Δn = 0.50, N > 20B-F, Δn = 0.50, N > 40B-F, Δn = 0.33, N > 20B-F, 0.70 ≤ ≤ 0.90
B-F, Δn = 0.50, N > 60B-F, = 1, N > 20

Conditions in which the F-statistic, adjusted F-tests, and B-F can be used for unbalanced designs with heterogeneity of variance and negative pairing.

G, group effect; T, time effect; G × T, group × time effect; Δn, coefficient of sample size variation; VR, variance ratio; γ1, skewness coefficient; γ2, kurtosis coefficient; N, total sample size; , Greenhouse-Geisser epsilon; F, F-statistic; F-GG, Greenhouse-Geisser adjusted F-test; F-HF, Huynh-Feldt adjusted F-test; B-F, bootstrap-F.

In an attempt to further simplify the interpretation of the results, Table 7 presents a concise synthesis of the recommendations derived from Tables 36. The table is intended as a practical guide for researchers and summarizes the procedures that performed adequately across the non-normal distributions examined, as well as under the different combinations of assumption violations investigated in the present study.

Table 7

Design typeEffectConditionRecommended test(s)
BalancedGGeneral recommendation under violation of assumptionsF, B-F
T, G × TModerate sphericity violation ( ≥ 0.80) with homogeneity and heterogeneityF
Moderate sphericity violation ( ≥ 0.70) with homogeneity and heterogeneityF-GG, F-HF
Severe sphericity violation ( = 0.50) with homogeneity and heterogeneityF-GG, F-HF (N > 60)
Severe sphericity violation ( 0.60) with homogeneity and heterogeneityF-GG, F-HF (N > 20)
General recommendation under violation of assumptionsB-F (N > 20)
Unbalanced, homogeneous variancesGGeneral recommendationF, and B-F (N > 20)
T, G × TModerate sphericity violation ( ≥ 0.70)F, F-GG, F-HF, B-F
Severe sphericity violation ( = 0.50)F-GG, F-HF (N > 60)
Severe sphericity violation ( = 0.60)F-GG, F-HF (N > 20)
Severe sphericity violation ( ≤ 0.60)B-F (N > 40)
Unbalanced, heterogeneous variances (positive pairing)GSmall inequality of group size (Δn = 0.16)F, B-F
Larger inequality of group size (Δn > 0.16)B-F
T, G × TSmall/moderate inequality of group size (Δn ≤ 0.33), moderate heterogeneity (VR ≤ 2), and moderate sphericity violation ( = 0.60-0.80)F-GG, F-HF (N > 20)
Small inequality of group size (Δn = 0.16), high heterogeneity (VR = 5), and severe sphericity violation ( ≤ 0.60)F-GG, F-HF (N > 20)
General recommendation under violation of assumptionsB-F (N > 20)
Unbalanced, heterogeneous variances (negative pairing)GSmall inequality of group size (Δn ≤ 0.16) and moderate heterogeneity (VR ≤ 2)F may be acceptable
Small and moderate inequality of group size (Δn ≤ 0.33) and moderate heterogeneity (VR ≤ 2)B-F
Larger inequality of group size (Δn = 0.50) and moderate/high heterogeneity (VR ≥ 2)B-F (N > 80)
T, G × TSmall inequality of group size (Δn ≤ 0.16), low heterogeneity (VR = 1.5), and sphericity satisfied ( ≈ 1)F may be acceptable
Small inequality of group size (Δn ≤ 0.16), low heterogeneity (VR = 1.5), and moderate sphericity violation ( ≥ 0.70)F-GG or F-HF may be acceptable
Small and moderate inequality of group size (Δn ≤ 0.33) and moderate heterogeneity (VR ≤ 2)B-F (N > 20)
Larger inequality of group size (Δn = 0.50) and high heterogeneity (VR ≥ 2)B-F (N > 60 - 80)

Practical recommendations for selecting robust tests in a 2 × 3 split-plot design across the non-normal distributions examined and other assumption violations (see Tables 36 for details).

G, group effect; T, time effect; G × T, group × time effect; Δn, coefficient of sample size variation; VR, variance ratio; N, total sample size; , Greenhouse-Geisser epsilon; F, F-statistic; F-GG, Greenhouse-Geisser adjusted F-test; F-HF, Huynh-Feldt adjusted F-test; B-F, bootstrap-F.

Finally, it should be highlighted that none of the procedures is robust under all conditions of non-sphericity, non-normality, heterogeneity of the variance, and sample size. All the statistics perform worse with severe simultaneous violations of all three assumptions when the group sizes are very unequal. The advice here is to try to plan the research with group sizes that are as similar as possible. In general, the bootstrap method represents an appropriate analytic alternative in most cases of violation of homogeneity of variances and sphericity, although it generally requires a sample size larger than 20 to control the Type I error rate in the majority of conditions. However, in some scenarios with negative pairing, it requires a larger sample (N > 80) to achieve robustness if the group sizes are very unequal.

This study has several limitations that should be noted. First, the results are applicable only to the conditions studied here, namely a design with 2 groups and 3 repeated measures, with skewness and kurtosis up to 2.31 and 8, respectively, with varying degrees of homogeneity and sphericity violations. Future studies could extend these results to other designs of higher order and other deviations from the underlying assumptions. Second, the findings are based solely on Type I error rates; future studies should complement these results by examining statistical power, which would provide a more comprehensive understanding of the performance of the procedures examined. Third, this study focused on a specific set of statistical procedures, and other approaches—such as linear mixed models, multivariate analysis of variance, permutation tests, robust statistics based on trimmed means—were not analyzed. Exploring these additional procedures in future research would help better handle repeated-measures designs in situations in which the underlying assumptions are violated.

Statements

Data availability statement

The datasets presented in this study can be found in the University of Malaga online repositoriy: https://hdl.handle.net/10630/38406. The datasets can also be found in the article/Supplementary material.

Author contributions

MJB: Writing – review & editing, Formal analysis, Writing – original draft, Data curation, Project administration, Conceptualization, Validation, Methodology, Supervision, Investigation. RA: Methodology, Investigation, Writing – review & editing, Writing – original draft, Visualization, Validation. RB: Resources, Validation, Investigation, Data curation, Supervision, Writing – review & editing, Conceptualization, Methodology, Software. JA: Validation, Writing – review & editing, Software. FJG-C: Validation, Writing – review & editing, Investigation, Data curation, Methodology, Visualization. GV: Software, Validation, Supervision, Writing – review & editing.

Funding

The author(s) declared that financial support was received for this work and/or its publication. This research was supported by grant PID2020-113191GB-I00, awarded through MCIN/AEI/10.13039/501100011033, and by the project PPRO-CTS110-G-2023 (CTS110-G-FEDER) funded by the University of Malaga through the Andalucía FEDER Programme 2021–2027.

Acknowledgments

The authors would like to thank Macarena Torrado for her collaboration in this study.

Conflict of interest

The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Generative AI statement

The author(s) declared that Generative AI was used in the creation of this manuscript. The author(s) used Microsoft Copilot exclusively during the preparation of the responses to reviewers for language revision and editing purposes. Microsoft Copilot was not used for data analysis, interpretation of results, methodological decisions, or manuscript content generation.

Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.

Publisher’s note

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.

Supplementary material

The Supplementary Material for this article can be found online at: https://www.frontiersin.org/articles/10.3389/fpsyg.2026.1787528/full#supplementary-material

References

Summary

Keywords

anova, bootstrap, Greenhouse-Geisser adjustment, Huynh-Feldt adjustment, repeated measures

Citation

Blanca MJ, Alarcón R, Bono R, Arnau J, García-Castro FJ and Vallejo G (2026) Guidance for data analysis in split-plot designs when parametric assumptions are violated: F-statistic, adjusted F-tests, and bootstrap-F. Front. Psychol. 17:1787528. doi: 10.3389/fpsyg.2026.1787528

Received

14 January 2026

Revised

10 July 2026

Accepted

13 July 2026

Published

06 August 2026

Volume

17 - 2026

Edited by

Fernando Marmolejo-Ramos, Flinders University, Australia

Reviewed by

Dolores Frias-Navarro, University of Valencia, Spain

AsIı Ceren Macunluoğlu, Independent Researcher, Düzce, Türkiye

Updates

Copyright

*Correspondence: María J. Blanca,

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Figures

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics