Abstract
Brain-imaging research has predominantly generated insight by means of classical statistics, including regression-type analyses and null-hypothesis testing using t-test and ANOVA. Throughout recent years, statistical learning methods enjoy increasing popularity especially for applications in rich and complex data, including cross-validated out-of-sample prediction using pattern classification and sparsity-inducing regression. This concept paper discusses the implications of inferential justifications and algorithmic methodologies in common data analysis scenarios in neuroimaging. It is retraced how classical statistics and statistical learning originated from different historical contexts, build on different theoretical foundations, make different assumptions, and evaluate different outcome metrics to permit differently nuanced conclusions. The present considerations should help reduce current confusion between model-driven classical hypothesis testing and data-driven learning algorithms for investigating the brain with imaging techniques.
“The trick to being a scientist is to be open to using a wide variety of tools.”
Breiman ()
Introduction
Among the greatest challenges humans face are cultural misunderstandings between individuals, groups, and institutions (Hall, 1976). The topic of the present paper is the culture clash between knowledge generation based on null-hypothesis testing and out-of-sample pattern generalization (Friedman, 1998; Breiman, ; Shmueli, 2010; Donoho, 2015). These statistical paradigms are now increasingly combined in brain-imaging studies (Kriegeskorte et al., 2009; Varoquaux and Thirion, 2014). Ensuing inter-cultural misunderstandings are unfortunate because the invention and application of new research methods has always been a driving force in the neurosciences (Greenwald, 2012; Yuste, 2015). Here the goal is to disentangle the contexts underlying classical statistical inference and out-of-sample generalization by providing a direct comparison of their historical trajectories, modeling philosophies, conceptual frameworks, and performance metrics.
During recent years, neuroscience has transitioned from qualitative reports of few patients with neurological brain lesions to quantitative lesion-symptom mapping on the voxel level in hundreds of patients (Gläscher et al., 2012). We have gone from manually staining and microscopically inspecting single brain slices to 3D models of neuroanatomy at micrometer scale (Amunts et al., ). We have also gone from experimental studies conducted by a single laboratory to automatized knowledge aggregation across thousands of previously isolated neuroimaging findings (Yarkoni et al., 2011; Fox et al., 2014). Rather than laboriously collecting in-house data published in a single paper, investigators are now routinely reanalyzing multi-modal data repositories (Derrfuss and Mar, 2009; Markram, 2012; Van Essen et al., 2012; Kandel et al., 2013; Poldrack and Gorgolewski, 2014). The detail of neuroimaging datasets is hence growing in terms of information resolution, sample size, and complexity of meta-information (Van Horn and Toga, 2014; Eickhoff et al., 2016; Bzdok and Yeo, ). As a consequence of the data demand of many pattern-recognition algorithms, the scope of neuroimaging analyses has expanded beyond the predominance of regression-type analyses combined with null-hypothesis testing (Figure 1). Applications of statistical learning methods (i) are more data-driven due to particularly flexible models, (ii) have scaling properties compatible with high-dimensional data with myriads of input variables, and (iii) follow a heuristic agenda by prioritizing useful approximations to patterns in data (Jordan and Mitchell, 2015; LeCun et al., 2015; Blei and Smyth, ). Statistical learning (Hastie et al., 2001) henceforth comprises the umbrella of “machine learning,” “data mining,” “pattern recognition,” “knowledge discovery,” “high-dimensional statistics,” and bears close relation to “data science.”
Figure 1
From a technical perspective, one should make a note of caution that holds across application domains such as neuroscience: While the research question often precedes the choice of statistical model, perhaps no single criterion exists that alone allows for a clear-cut distinction between classical statistics and statistical learning in all cases. For decades, the two statistical cultures have evolved in partly independent sociological niches (Breiman,
More generally, ensuring that a statistical effect discovered in one set of data extrapolates to new observations in the brain can take different forms (Efron, 2012). As one possible definition, “the goal of statistical inference is to say what we have learned about the population X from the observed data x” (Efron and Tibshirani, 1994). In a similar spirit, a committee report to the National Academies of the USA stated (Committee on the Analysis of Massive Data et al.,
For a long time, knowledge generation in psychology, neuroscience, and medicine has been dominated by classical statistics with estimation of linear-regression-like models and subsequent statistical significance testing whether an effect exists in the sample. In contrast, computation-intensive pattern learning methods have always had a strong focus on prediction in frequently extensive data with more modest concern for interpretability and the “right” underlying question (Hastie et al., 2001; Ghahramani, 2015). In many statistical learning applications, it is standard practice to quantify the ability of a predictive pattern to extrapolate to other samples, possibly in individual subjects. In a two-step procedure, a learning algorithm is fitted on a typically bigger amount of available data (training data) and the ensuing fitted model is empirically evaluated on a commonly smaller amount of independent data (test data). This stands in contrast to classical statistical inference where the investigator seeks to reject the null hypothesis by considering the entirety of a data sample (Wasserstein and Lazar, 2016), typically all available subjects. In this case, the desired relevance of a statistical relationship in the underlying population is ensured by formal mathematical proofs and is not commonly ascertained by explicit evaluations on new data (Breiman,
Taking an epistemological perspective helps appreciating that scientific research is rarely an entirely objective process but deeply depends on the beliefs and expectations of the investigator. A new “scientific fact” about the brain is probably not established in vacuo (Fleck et al., 1935; terms in quotes taken from source). Rather, a research “object” is recognized and accepted by the “subject” according to socially conditioned “thought styles” that are cultivated among members of “thought collectives.” A witnessed and measured neurobiological phenomenon tends to only become “true” if not at odds with the constructed “thought history” and “closed opinion system” shared by that subject. The present paper will revisit and reintegrate two such thought milieus in the context of imaging neuroscience: classical statistics (ClSt) and statistical learning (StLe).
Different histories: the origins of classical hypothesis testing and pattern-learning algorithms
One of many possible ways to group statistical methods is by framing them along the lines of ClSt and StLe. The incongruent historical developments of the two statistical communities are even evident from their basic terminology. Inputs to statistical models are usually called independent variables, explanatory variables, or predictors in the ClSt community, but are typically called features collected in a feature space in the StLe community. The model outputs are typically called dependent variables, explained variable, or responses in ClSt, while these are often called target variables in StLe. It follows a summary of characteristic events in the development of what can today be considered as ClSt and StLe (Figure 2).
Figure 2

Developments in the history of classical statistics and statistical learning. Examples of important inventions in statistical methodology. Roughly, a number of statistical methods taught in today's textbooks in psychology and medicine have emerged in the first half of the twentieth century (blue). Instead, many algorithmic techniques and procedures have emerged in the second half of the twentieth century (red). “The postwar era witnessed a massive expansion of statistical methodology, responding to the data-driven demands of modern scientific technology.” (Efron and Hastie, 2016).
Around 1900 the notions of standard deviation, goodness of fit, and the p < 0.05 threshold emerged (Cowles and Davis,
It is a topic of current debate1,2,3 whether ClSt is a discipline that is separate from StLe (e.g., Chambers,
As often cited beginnings of statistical learning approaches, the perceptron was an early brain-inspired computing algorithm (Rosenblatt, 1958), and Arthur Samuel created a checker board program that succeeded in beating its own creator (Samuel, 1959). Such studies toward artificial intelligence (AI) led to enthusiastic optimism and subsequent periodes of disappointment during the so-called “AI winters” in the late 70s and around the 90s (Russell and Norvig, 2002; Kurzweil, 2005; Cox and Dean,
In sum, “the biggest difference between pre- and post-war statistical practice is the degree of automation” (Efron and Tibshirani, 1994) up to a point where “almost all topics in twenty-first-century statistics are now computer-dependent” (Efron and Hastie, 2016). ClSt has seen many important inventions in the first half of the twentieth century, which have often developed at statistical departments of academic institutions and remain in nearly unchanged form in current textbooks of psychology and other empirical sciences. The emergence of StLe as a coherent field has mostly taken place in the second half of the twentieth century as a number of disjoint developments in industry and often non-statistical departments in academia (e.g., AT&T Bell Laboratories), which lead for instance to artificial neural networks, support vector machines, and boosting algorithms (Efron and Hastie, 2016). Today, systematic education in StLe is still rare at the large majority of universities, in contrast to the many consistently offered ClSt courses (Cleveland,
In neuroscience, the advent of brain-imaging techniques, including positron emission tomography (PET) and functional magnetic resonance imaging (fMRI), allowed for the in-vivo characterization of the neural correlates underlying sensory, cognitive, or affective tasks. Brain scanning enabled quantitative brain measurements with many variables per observation (analogous to the advent of high-dimensional microarrays in genetics; Efron, 2012). Since the inception of PET and fMRI, deriving topographical localization of neural activity changes was dominated by analysis approaches from ClSt, especially the general linear model (Scheffé, 1959; Poline and Brett, 2012; GLM). The classical approach to neuroimaging analysis is probably best exemplified by the statistical parametric mapping (SPM) software package that implements the GLM to provide a mass-univariate characterization of regionally specific effects.
As distributed information over voxels is less well captured by many ClSt approaches, including common GLM applications, StLe models were proposed early on for neuroimaging investigations. For instance, principal component analysis was used to distinguish globally distributed neural activity changes (Moeller et al., 1987) as well as to study Alzheimer's disease (Grady et al., 1990). Canonical correlation analysis was used to quantify complex relationships between task-free neural activity and schizophrenia symptoms (Friston et al., 1992). However, these first approaches to “multivariate” brain-behavior associations did not ignite a major research trend (cf. Worsley et al., 1997; Friston et al., 2008). As a seminal contribution, Haxby and colleagues devised an innovative across-voxel correlation analysis to provide evidence against the widely assumed face-specificity of neural responses in the ventral temporal cortex (2001). This ClSt realization of one-nearest neighbor classification based on correlation distance foreshadowed several important developments, including (i) joint analysis of sets of brain locations to capture “distributed and overlapping representations”, (ii) repeated analysis in different splits of the data sample to compare against chance performance, and (iii) analysis across multiple stimulus categories to assess the specificity of neural responses. The finding of distributed face representation was confirmed in independent, similar data (Cox and Savoy,
The application of StLe methods in neuroimaging increased further after rebranding as “mind-reading,” “brain decoding,” and “MVPA” (Haynes and Rees, 2005; Kamitani and Tong, 2005). Note that “MVPA” initally referred to “multivoxel pattern analysis” (Kamitani and Tong, 2005; Norman et al., 2006) and later changed to “multivariate pattern analysis” (Haynes and Rees, 2005; Hanke et al., 2009; Haxby, 2012). Up to that point, the term prediction had less often been used by imaging neuroscientists in the sense of out-of-sample generalization of a learning algorithm and more often in the incompatible sense of (in-sample) linear correlation such as using Pearson's or Spearman's method (Shmueli, 2010; Gabrieli et al., 2015). While there was scarce discussion of the position of “decoding” models in formal statistical terms, growing interest was manifested in first review publications and tutorial papers on applying StLe methods to neuroimaging data (Haynes and Rees, 2006; Mur et al., 2009; Pereira et al., 2009). The interpretational gains of this new access to the neural representation of behavior and its disturbances in disease was flanked by the availability of necessary computing power and memory resources. Although challenging to realize, “deep” neural network algorithms have recently been introduced to neuroimaging research (Plis et al., 2014; de Brebisson and Montana, 2015; Güçlü and van Gerven, 2015). These computation-intensive models might help in approximating and deciphering the nature of neural processing in brain circuits (Cox and Dean,
From a conceptual viewpoint (Figure 3), a large majority of statistical methods can be situated somewhere on a continuum between the two poles of ClSt and StLe (Committee on the Analysis of Massive Data et al.,
Figure 3

Key differences in the modeling philosophy of classical statistics and statistical learning. Ten modeling intuitions that tend to be relatively more characteristic for classical statistical methods (blue) or pattern-learning methods (red). In comparison to ClSt, StLe “is essentially a form of applied statistics with increased emphasis on the use of computers to statistically estimate complicated functions and a decreased emphasis on proving confidence intervals around these functions” (Goodfellow et al., 2016). Broadly, ClSt tends to be more analytical by imposing mathematical rigor on the phenomenon, whereas StLe tends to be more heuristic by finding useful approximations. In practice, ClSt is probably more often applied to experimental data, where a set of target variables are systematically controlled by the investigator and the brain system under studied has been subject to experimental perturbation. Instead, StLe is probably more often applied to observational data without such structured influence and where the studied system has been left unperturbed. ClSt fully specifies the statistical model at the beginning of the investigation, whereas in StLe there is a bigger emphasis on models that can flexibly adapt to the data (e.g., learning algorithms creating decision trees).
Case study one: cognitive contrast analysis and decoding mental states
Vignette: A neuroimaging investigator wants to reveal the neural correlates underlying face processing in humans. 40 healthy, right-handed adults are recruited and undergo a block design experiment run in a 3T MRI scanner with whole-brain coverage. In a passive viewing paradigm, 60 colored stimuli of unfamiliar faces are presented, which have forward head and gaze position. The control condition presents colored pictures of 60 different houses to the participants. In the experimental paradigm, a picture of a face or a house is presented for 2 s in each trial and the interval between trials within each block is randomly jittered varying from 2 to 7 s. The picture stimuli are presented in pseudo-randomized fashion and are counterbalanced in each passively watching participant. Despite the blocked presentation of stimuli, each experiment trial is modeled separately. The fMRI data are analyzed using a GLM as implemented in the SPM software package. Two task regressors are included in the model for the face and house conditions based on the stimulus onsets and viewing durations and using a canonical hemodynamic response function. In the GLM design matrix, the face column and house column are hence set to 1 for brain scans from the corresponding task condition and set to 0 otherwise. Separately in each brain voxel, the GLM parameters are estimated, which fits betaface and betahouse regression coefficients to explain the contribution of each experimental task to the neural activity increases and decreases observed in that voxel. A t-test can then formally assess whether the fMRI signal in the current voxel is significantly more involved in viewing faces as opposed to the house control condition.
Question: What is the statistical difference between subtracting the neural activity from the face vs. house conditions and decoding the neural activity during face vs. house processing?
Computing cognitive contrasts is a ClSt approach that was and still is routinely performed in the mass-univariate regime: it fits a separate GLM model for each individual voxel in the brain scans and then tests for significant differences between the obtained condition coefficients (Friston et al., 1994). Instead, decoding cognitive processes from neural activity is a StLe approach that is typically performed in a multivariate regime: a learning algorithm is trained on a large number of voxel observations in brain scans and then the model's prediction accuracy is evaluated on sets of new brain scans. These ClSt and StLe approaches to identifying the neural correlates underlying cognitive processes of interest are closely related to the notions of encoding models and decoding models, respectively (Kriegeskorte, 2011; Naselaris et al., 2011; Pedregosa et al., 2015; but see Güçlü and van Gerven, 2015).
Encoding models regress the brain data against a design matrix with indicators of the face vs. house condition and formally test whether the difference is statistically significant. Decoding models typically aim to predict these indicators by training and empirically evaluating classification algorithms on different splits from the whole dataset. In ClSt parlance, the model explains the neural activity, the dependent or explained variable, measured in each separate brain voxel, by the beta coefficients according to the experimental condition indicators in the design matrix columns, the independent or explanatory variables. That is, the GLM can be used to explain neural activity changes by a linear combination of experimental variables (Naselaris et al., 2011). Answering the same neuroscientific question with decoding models in StLe jargon, the model weights of a classifier are fitted on the training set of the input data to predict the class labels, the target variables, and are subsequently evaluated on the test set by cross-validation to obtain their out-of-sample generalization performance. Here, classification algorithms are used to predict entries of the design matrix by identifying a linear or more complicated combination between the many simultaneously considered brain voxels (Pereira et al., 2009). More broadly, ClSt applications in functional neuroimaging tend to estimate the location of cognitive processes from neural activity, whereas many StLe applications estimate properties of neural activity underlying different cognitive tasks.
A key difference between many ClSt-mediated encoding models and StLe-mediated decoding models thus pertains to the direction of statistical estimation between brain space and behavior space (Friston et al., 2008; Varoquaux and Thirion, 2014). It was noted (Friston et al., 2008) that the direction of brain-behavior association is related to the question whether the stimulus indicators in the model act as causes by representing deterministic experimental variables of an encoding model or consequences by representing probabilistic outputs of a decoding model. Such considerations also reveal the intimate relationship of ClSt models to the notion of forward inference, while StLe methods are probably more often used for formal reverse inference in functional neuroimaging (Poldrack, 2006; Eickhoff et al., 2011; Yarkoni et al., 2011; Varoquaux and Thirion, 2014). On the one hand, forward inference relates to encoding models by testing the probability of observing activity in a brain location given knowledge of a psychological process. On the other hand, reverse inference relates to brain decoding to the extent that classification algorithms can learn to distinguish experimental fMRI data to belong to two psychological conditions and subsequently be used to estimate the presence of specific cognitive processes based on new neural activity observations (cf. Poldrack, 2006). Finally, establishing a brain-behavior association has been argued to be more important than the actual direction of the mapping function (Friston, 2009). This author stated that “showing that one can decode activity in the visual cortex to classify […] a subject's percept is exactly the same as demonstrating significant visual cortex responses to perceptual changes” and, conversely, “all demonstrations of functionally specialized responses represent an implicit mindreading.”
Conceptually, GLM-based encoding models follow a localization agenda by testing hypotheses on regional effects of functional specialization in the brain (where?). A t-test is used to compare pairs of neural activity estimates to statistically distinguish the target face and the non-target house condition (Friston et al., 1996). Essentially, this test for significant differences between the fitted beta coefficients corresponds to two stimulus indicators based on well-founded arguments from cognitive theory. This statistical approach assumes that cognitive subtraction is possible, that is, the regional brain responses of interest can be isolated by contrasting two sets of brain scans that are believed to differ in the cognitive facet of interest (Friston et al., 1996; Stark and Squire, 2001). For one voxel location at a time, an attempt is made to reject the null hypothesis of no difference between the averaged neural activity level of a target brain state and the averaged neural activity of a control brain state. It is important to appreciate that the localization agenda thus emphasizes the relative difference in fMRI signal during tasks and may neglect the individual neural activity information of each particular task (Logothetis et al., 2001). Note that the univariate GLM analysis can be extended to more than one output (dependent or explained) variable within the ClSt regime by performing a multivariate analysis of covariance (MANCOVA). This allows for tests of more complex hypotheses but incurs multivariate normality assumptions (Kriegeskorte, 2011).
More generally, it is seldom mentioned that the standard GLM would not have been solvable for unique solutions in the high-dimensional “n < < p” regime, instead of fitting one model for each voxel in the brain scans. This is because the number of brain voxels p exceed by far the number of data samples n (i.e., leading to an under-determined system of equations), which incapacitates many statistical estimators from ClSt (cf. Giraud, 2014; Hastie et al., 2015). Regularization by sparsity-inducing norms, such as in modern penalized regression analysis using the LASSO and ElasticNet, emerged only later (Tibshirani, 1996; Zou and Hastie, 2005) as a principled StLe strategy to de-escalate the need for dimensionality reduction or preliminary filtering of important voxels and to enable the tractability of the high-dimensional analysis setting.
Because hypothesis testing for significant differences between beta coefficients of fitted GLMs relies on comparing the means of neural activity measurements, the results from statistical tests are not corrupted by the conventionally applied spatial smoothing with a Gaussian filter. On the contrary, this image preprocessing step even helps the correction for multiple comparisons based on random fields theory (cf. below), alleviates inter-individual neuroanatomical variability, and can thus increases sensitivity. Spatial smoothing however discards fine-grained neural activity patterns spatially distributed across voxels that potentially carry information associated with mental operations (cf. Kamitani and Sawahata, 2010; Haynes, 2015). Indeed, some authors believe that sensory, cognitive, and motor processes manifest themselves as “neuronal population codes” (Averbeck et al.,
In so doing, decoding models use learning algorithms in an information agenda by showing generalization of robust patterns to new brain activity acquisitions (Kriegeskorte et al., 2006; Mur et al., 2009; de-Wit et al., 2016). Information that is weak in one voxel but spatially distributed across voxels can be effectively harvested in a structure-preserving fashion (Haynes and Rees, 2006; Haynes, 2015). This modeling agenda is focused on the whole neural activity pattern, in contrast to the localization agenda dedicated to separate increases or decreases in neural activity level. For instance, the default mode network typically exhibits activity decreases at the onset of many psychological tasks with visual or other sensory stimuli, whereas the induced activity patterns in that less activated network may nevertheless functionally subserve task execution (Bzdok et al.,
Among other views, it has previously been proposed (Brodersen,
In sum, the statistical properties of ClSt and StLe methods have characteristic consequences in neuroimaging analysis and interpretation. They can hence offer different access routes and complementary answers to identical neuroscientific questions.
Case study two: small volume correction and searchlight analysis
Vignette: The neuroimaging experiment from case study 1 successfully identified the fusiform gyrus of the ventral visual stream to be more responsive to face stimuli than house stimuli. However, the investigator's initial hypothesis of also observing face-responsive neural activity in the ventromedial prefrontal cortex could not be confirmed in the whole-brain analyses. The investigator therefore wants to follow up with a topographically focused approach that examines differences in neural activity between the face and house conditions exclusively in the ventromedial prefrontal cortex.
Question: What are the statistical implications of delineating task-relevant neural responses in a spatially constrained search space rather than analyzing brain measurements of the entire brain?
A popular ClSt approach to corroborate less pronounced neural activity findings is small volume correction. This region of interest (ROI) analysis involves application of the mass-univariate GLM approach only to the ventromedial prefrontal cortex as a preselected biological compartment, rather than considering the gray-matter voxels of the entire brain in a naïve, topographically unconstrained fashion. Small volume correction allows for significant findings in the ROI that remain sub-threshold after accounting for the tens of thousands of multiple comparisons in the whole-brain GLM analysis. Small volume correction is therefore a simple means to alleviate the multiple-comparisons problem that motivated more than two decades of still ongoing methodological developments in the neuroimaging domain (Worsley et al., 1992; Smith et al., 2001; Friston, 2006; Nichols, 2012). Whole-brain GLM results were initially reported as uncorrected findings without accounting for multiple comparisons, then with Bonferroni's family wise error (FWE) correction, later by random field theory correction using neural activity height (or clusters), followed by false discovery rate (FDR) (Genovese et al., 2002) and slowly increasing adoption of cluster-thresholding for voxel-level inference via permutation testing (Smith and Nichols, 2009). Rather than the isolated voxel, it has early been discussed that a possibly better unit of interest should be spatially neighboring voxel groups (see here for an overview: Chumbley and Friston,
A related cousin of small volume correction in the StLe world would be to apply classification algorithms to a subset of voxels to be considered as input to the model (i.e., feature selection). In particular, searchlight analysis is an increasingly popular learning technique that can identify locally constrained multivariate patterns in neural activity (Friman et al., 2001; Kriegeskorte et al., 2006). For each voxel in the ventromedial prefrontal cortex, the brain measurements of the immediate neighborhood are first collected (e.g., radius of 10 mm voxels). In each such searchlight, a classification algorithm, for instance linear support vector machines, is then trained on one part of the brain scans (training set) and subsequently applied to determine the prediction accuracy in the remaining, unseen brain scans (test set). In this StLe approach, the excess of brain voxels is handled by performing pattern recognition analysis in only dozens of locally adjacent voxel neighborhoods at a time. Finally, the mean classification accuracy of face vs. house stimuli across all permutations over the brain data is mapped to the center of each considered sphere. The searchlight is then moved through the ROI until each seed voxel had once been the center voxel of the searchlight. This yields a voxel-wise classification map of accuracy estimates for the entire ventromedial prefrontal cortex. Consistent with the information agenda (cf. above), searchlight analysis quantifies the extent to which (local) neural activity patterns can predict the difference between the house and face conditions. It contrasts small volume correction that determines whether one experimental condition exhibited a significant neural activity increase or decrease relative to a particular other experimental condition, consistent with the localization agenda. Further, searchlight analysis alleviates the burden of abundant input variables by fitting learning algorithms restricted to the voxels in small sphere neighborhoods. However, the searchlight procedure thus yields many prediction performances for many brain locations, which motivates correction for multiple comparisons across the considered neighborhoods.
When considering high-dimensional brain scans through the ClSt lens, the statistical challenge resides in solving the multiple-comparisons problem (Nichols and Hayasaka, 2003; Nichols, 2012). From the StLe stance, however, it is the curse of dimensionality and overfitting that statistical analyses need to tackle (Friston et al., 2008; Domingos, 2012). Many neuroimaging analyses based on ClSt methods can be viewed as testing a particular hypothesis (i.e., the null hypothesis) repeatedly in a large number of separate voxels. In contrast, testing whether learning algorithm extrapolate to new brain data can be viewed as searching through thousands of different hypotheses in a single process (i.e., walking through the hypothesis space; cf. above) (Shalev-Shwartz and Ben-David, 2014).
As common brain scans offer measurements of >100,000 brain locations, a mass-univariate GLM analysis typically entails the same statistical test to be applied >100,000 times. The more often the investigator tests a hypothesis of relevance for a brain location, the more locations will be falsely detected as relevant (false positive, Type I error), especially in the noisy neuroimaging data. All dimensions in the brain data (i.e., voxel variables) are implicitly treated as equally important and no neighborhoods of most expected variation are statistically exploited (Hastie et al., 2001). Hence, the absence of restrictions on observable structure in the set of data variables during the statistical modeling of neuroimaging data takes a heavy toll at the final inference step. This is where random field theory comes to the rescue. As noted above, this form of topological inference dispenses with the problem of inferring which voxels are significant and tries to identify significant topological features in the underlying distributed responses. By definition, topological features like maxima are sparse events and can be thought of as a form of dimensionality reduction—not in data space but in the statistical characterization of where neural responses occur.
This is contrasted by the high-dimensional StLe regime, where the initial model family chosen by the investigator determines the complexity restrictions to all data dimensions (i.e., all voxels, not single voxels) that are imposed explicitly or implicitly by the model structure. Model choice predisposes existing but unknown low-dimensional neighborhoods in the full voxel space to achieve the prediction task. Here, the toll is taken at the beginning of the investigation because there are so many different alternative model choices that would impose a different set of complexity constraints to the high-dimensional measurements in the brain. For instance, signals from “brain regions” are likely to be well approximated by models that impose discrete, locally constant compartments on the data (e.g., k-means or spatially constrained Ward clustering). Instead, tuning model choice to signals from macroscopical “brain networks” should impose overlapping, locally continuous data compartments (e.g., independent component analysis or sparse principal component analysis) (Yeo et al., 2014; Bzdok and Yeo,
Exploiting such effective dimensions in the neuroimaging data (i.e., coherent brain-behavior associations involving many distributed brain voxels) is a rare opportunity to simultaneously reduce the model bias and model variance, despite their typical inverse relationship (Hastie et al., 2001). Model bias relates to prediction failures incurred because the learning algorithm can systematically not represent certain parts of the underlying relationship between brain scans and experimental conditions (formally, the deviation between the target function and the average function space of the model). Model variance relates to prediction failures incurred by noise in the estimation of the optimal brain-behavior association (formally, the difference between the best-choice input-output relation and the average function space of the model). A model that is too simple to capture a brain-behavior association probably underfits due to high bias. Yet, an overly complex model probably overfits due to high variance. Generally, high-variance approaches are better at approximating the “true” brain-behavior relation (i.e., in-sample model estimation), while high-bias approaches have a higher chance of generalizing the identified pattern to new observations (i.e., out-of-sample model evaluation). The bias-variance tradeoff can be useful in explaining why applications of statistical models intimately depend on (i) the amount of available data, (ii) the typically not known amount of noise in the data, and (iii) the unknown complexity of the target function in nature (Abu-Mostafa et al.,
Learning algorithms that overcome the curse of dimensionality—extracting coherent patterns from all considered brain voxels at once—typically incorporate an implicit bias for anisotropic neighborhoods in the data (Hastie et al., 2001; Bach,
As an practical summary, drawing classical inference in neuroimaging data has largely been performed by considering each voxel independently and by massive simultaneous testing of a same null hypothesis in all observed voxels. This has incurred a multiple-comparisons problem difficult enough that common approaches may still be prone to incorrect results (Efron, 2012). In contrast, aiming for generalization of a pattern in high-dimensional neuroimaging data to new observations in the brain incurs the equally challenging curse of dimensionality. Successfully accounting for the high number of input dimensions will probably depend on learning models that impose neurobiologically justified bias and keeping the variance under control by dimensionality reduction and regularization techniques.
More broadly, asking at what point new neurobiological knowledge is arising during ClSt and StLe investigations relies on largely distinct theoretical frameworks that revolve around null-hypothesis testing and statistical learning theory (Figure 4). Both ClSt and StLe methods share the common goal of demonstrating relevance of a given effect in the data beyond the sample brain scans at hand. However, the attempt to show successful extrapolation of a statistical relationship at the general population is embedded in different mathematical contexts. Knowledge generation in ClSt and StLe is hence rooted in different notions of statistical inference.
Figure 4

Key concepts in classical statistics and statistical learning. Schematic with statistical notions that are relatively more associated with classical statistical methods (left column) or pattern-learning methods (right column). As there is a smooth transition between the classical statistical toolkit and learning algorithms, some notions may be closely associated with both statistical cultures (middle column).
ClSt laid down its most important inferential framework in the Popperian spirit of critical empiricism (Popper, 1935/2005): scientific progress is to be made by continuous replacement of current hypotheses by ever more pertinent hypotheses using falsification. The rationale behind hypothesis falsification is that one counterexample can reject a theory by deductive reasoning, while any quantity of evidence can not confirm a given theory by inductive reasoning (Goodman, 1999). The investigator verbalizes two mutually exclusive hypotheses by domain-informed judgment. The alternative hypothesis should be conceived as the outcome intended by the investigator and to contradict the state of the art of the research topic. The null hypothesis represents the devil's advocate argument that the investigator wants to reject (i.e., falsify) and it should automatically deduce from the newly articulated alternative hypothesis. A conventional 5%-threshold (i.e., equating with roughly two standard deviations) guards against rejection due to the idiosyncrasies of the sample that are not representative of the general population. If the data have a probability of ≤5% given the null hypothesis [P(result|H0)], it is evaluated to be significant. Such a test for statistical significance indicates a difference between two means with a 5% chance of being a false positive finding. If the null hypothesis can not be rejected (which depends on power), then the test yields no conclusive result, rather than a null result (Schmidt, 1996). In this way, classical hypothesis testing continuously replaces currently embraced hypotheses explaining a phenomenon in nature by better hypotheses with more empirical support in a Darwinian selection process. Finally, Fisher, Neyman, and Pearson intended hypothesis testing as a marker for further investigation, rather than an off-the-shelf decision-making instrument (Cohen,
In StLe instead, answers to how neurobiological conclusions can be drawn from a dataset at hand are provided by the Vapnik-Chervonenkis dimensions (VC dimensions) from statistical learning theory (Vapnik, 1989, 1996). The VC dimensions of a pattern-learning algorithm quantify the probability at which the distinction between the neural correlates underlying the face vs. house conditions can be captured and used for correct predictions in new, possibly later acquired brain scans from the same cognitive experiment (i.e., out-of-sample generalization). Such statistical approaches implement the inductive strategy to learn general principles (i.e., the neural signature associated with given cognitive processes) from a series of exemplary brain measurements, which contrasts the deductive strategy of rejecting a certain null hypothesis based on counterexamples (cf. Tenenbaum et al., 2011; Bengio,
As one of the most important results from statistical learning theory, in any intelligent learning system, the opportunity to derive abstract patterns in the world by reducing the discrepancy between prediction error from training data (in-sample estimate) and prediction error from independent test data (out-of-sample estimate) decreases with the higher model capacity and increases with the number of available training observations (Vapnik and Kotz, 1982; Vapnik, 1996). In brain imaging, a learning algorithm is hence theoretically backed up to successfully predict outcomes in future brain scans with high probability if the choosen model ignores structure that is overly complicated, such as higher-order non-linearities between many brain voxels, and if the model is provided with a sufficient number of training brain scans. Hence, VC dimensions provide explanations why increasing the number of considered brain voxels as input features (i.e., entailing increased number of model parameters) or using a more sophisticated prediction model, requires more training data for successful generalization. Notably, the VC dimensions (analogous to null-hypothesis testing) are unrelated to the target function, as the “true” mechanisms underlying the studied phenomenon in nature. Nevertheless, the VC dimensions provide justification that a certain learning model can be used to approximate that target function by fitting a model to a collection of input-output pairs. In short, VC dimensions is among the best frameworks to derive theoretical errors bounds for predictive models (Abu-Mostafa et al.,
Further, some common invalidations of the ClSt and StLe statistical concern in neuroimaging studies performing classical inference is double dipping or circular analysis (Kriegeskorte et al., 2009). This occurs when, for instance, first correlating a behavioral measure with brain activity and then using the identified subset of brain voxels for a second correlation analysis with that same behavioral measurement (Lieberman et al., 2009; Vul et al., 2009). In this scenario, voxels are submitted to two statistical tests with the same goal in a nested, non-independent fashion5 (Freedman, 1983). This corrupts the validity of the null hypothesis on which the reported test results conditionally depend. Importantly, this case of repeating a same statistical estimation with iteratively pruned data selections (on the training data split) is a valid routine in the StLe framework, such as in recursive feature extraction (Guyon et al., 2002; Hanson and Halchenko, 2008). However, double-dipping or circular analysis in ClSt applications to neuroimaging data have an analog in StLe analyses aiming at out-of-sample generalization: data-snooping or peeking (Pereira et al., 2009; Abu-Mostafa et al.,
In sum, statistical inference in ClSt is drawn by using the entire data at hand to formally test for theoretically guaranteed extrapolation of an effect to the general population. In stark contrast, inferential conclusions in StLe are typically drawn by fitting a model on a larger part of the data at hand (i.e., in-sample model selection) and empirically testing for successful extrapolation to an independent, smaller part of the data (i.e., out-of-sample model evaluation). As such, ClSt has a focus on in-sample estimates and explained-variance metrics that measure some form of goodness of fit, while StLe has a focus on out-of-sample estimates and prediction accuracy.
Case study three: significant group differences and predicting the group of participants
Vignette: After isolating the neural correlates underlying face processing, the neuroimaging investigator wants to examine their relevance in psychiatric disease. In addition to the 40 healthy participants, 40 patients diagnosed with schizophrenia are recruited and administered the same experimental paradigm and set of face and house pictures. In this clinical fMRI study on group differences, the investigator wants to explore possible imaging-derived markers that index deficits in social-affective processing in patients carrying a diagnosis of schizophrenia.
Question: Can metrics of statistical relevance from ClSt and StLe be combined to corroborate a given candidate biomarker?
Many investigators in imaging neuroscience share a background in psychology, biology, or medicine, which includes training in traditional “textbook” statistics. Many neuroscientists have thus adopted a natural habit of assessing the quality of statistical relationships by means of p-values, effect sizes, confidence intervals, and statistical power. These are ubiquitously taught and used at many universities, although they are not the only coherent set of statistical diagnostics (Figure 5). These outcome metrics from ClSt may for instance be less familiar to some scientists with a background in computer science, physics, engineering, or philosophy. As an equally legitimate and internally coherent, yet less widely known diagnostic toolkit from the StLe community, prediction accuracy, precision, recall, confusion matrices, F1 score, and learning curves can also be used to measure the relevance of statistical relationships (Abu-Mostafa et al.,
Figure 5

Key differences between measuring outcomes in classical statistics and statistical learning. Ten intuitions on quantifying statistical modeling outcomes that tend to be relatively more true for classical statistical methods (blue) or pattern-learning methods (red). ClSt typically yields point estimates and interval estimates (e.g., p-values, variances, confidence intervals), whereas StLe frequently outputs a function or a program that can yield point and interval estimates on new observations (e.g., the k-means centroids or a trained classifier's decision function can be applied to new data). In many cases, classical inference is a judgment about an entire data sample, whereas a trained predictive model can obtain quantitative answers from a single data point.
On a general basis, applications of ClSt and StLe methods may not judge findings on identical grounds (Breiman,
For neuroscientists adopting a ClSt culture computing p-values takes a central position. The p-value denotes the probability of observing a result at least as extreme as a test statistic, assuming the null hypothesis is true. Results are considered significant when it is equal or below a pre-specified value, like p = 0.05 (Anderson et al.,
The essentially binary p-value (i.e., significant vs. not significant) is therefore often complemented by continuous effect size measures for the importance of rejecting H0. The effect size allows the identification of marginal effects that pass the statistical significance threshold but are not practically relevant in the real world. The p-value is a deductive inferential measure, whereas the effect size is a descriptive measure that follows neither inductive nor deductive reasoning. The (normalized) effect size can be viewed as the strength of a statistical relationship—how much H0 deviates from H1, or the likely presence of an effect in the general population (Chow,
Additionally, the certainty of a point estimate (i.e., the outcome is a value) can be expressed by an interval estimate (i.e., the outcome is a value range) using confidence intervals (Casella and Berger,
Some confidence intervals can be computed in various data scenarios and statistical regimes, whereas the power may be especially meaningful within the culture of classical hypothesis testing (Cohen,
While neuroimaging studies based on classical statistical inference ubiquitously report p-values and confidence intervals, there have however been few reports of effect size in the neuroimaging literature (Kriegeskorte et al., 2010). Effect sizes are however necessary to compute power estimates. This explains the even rarer occurrence of power calculations in the neuroimaging literature (Yarkoni and Braver, 2010; but see Poldrack et al., 2017). Given the importance of p-values and effect sizes, the goal of computing both these useful statistics, such as for group differences in the neural processing of face stimuli, can be achieved based on two independent samples of these experimental data (especially if some selection process has been used). One sample would be used to perform statistical inference on the neural activity change yielding a p-value and one sample to obtain unbiased effect sizes. Further, it has been previously emphasized (Friston, 2012) that p-values and effect sizes reflect in-sample estimates in a retrospective inference regime (ClSt). These metrics find an analog in out-of-sample estimates issued from cross-validation in a prospective prediction regime (StLe). In-sample effect sizes are typically an optimistic estimate of the “true” effect size (inflated by high significance thresholds), whereas out-of-sample effect sizes are unbiased estimates of the “true” effect size.
In the high-dimensional scenario, the StLe-minded investigator analyzing “wide” neuroimaging data in our case, computing, and judging statistical significance by p-values can become challenging (Bühlmann and Van De Geer,
The neuroscientist who adopted a StLe culture is in the habit of corroborating prediction accuracies using cross-validation: the de facto standard to obtain an unbiased estimate of a model's capacity to generalize beyond the brain scans at hand (Hastie et al., 2001; Bishop,
The classification accuracy can be further decomposed into group-wise metrics based on the so-called confusion matrix, the juxtaposition of the true and predicted group memberships. The precision measures (Table 1) how many of the labels predicted from brain scans are correct, that is, how many participants predicted to belong to a certain class really belong to that class. Put differently, among the participants predicted to suffer from schizophrenia, how many have really been diagnosed with that disease? On the other hand, the recall measures how many labels are correctly predicted, that is, how many members of a class were predicted to really belong to that class. Hence, among the participants known to be affected by schizophrenia, how many were actually detected as such? Precision can be viewed as a measure of “exactness” and recall as a measure of “completeness” (Powers, 2011).
Table 1
| Notion | Formula |
|---|---|
| Specificity | true negative/(true negative + false positive) |
| Sensitivity/Recall | true positive/(true positive + false negative) |
| Precision | true positive/(true positive + false positive) |
Metrics used to create ROC curves.
Neither accuracy, precision, or recall allow injecting subjective importance into the evaluation process of the learning algorithm. This disadvantage is addressed by the Fbetascore: a weighted combination of the precision and recall prediction scores. Concretely, the F1 score would equally weigh precision and recall of class predictions, while the F0.5 score puts more emphasis on precision and the F2 score more on recall. Moreover, applications of recall, precision, and Fbeta scores have been noted to ignore the true negative cases as well as to be highly susceptible to estimator bias (Powers, 2011). Needless to say, no single outcome metric can be equally optimal in all contexts.
Extending from the setting of healthy-diseased classification to the multi-class setting (e.g., comparing healthy, schizophrenic, bipolar, and autistic participants) injects ambiguity into the interpretation of accuracy scores. Rather than reporting mere better-than-chance findings in StLe analyses, it becomes more important to evaluate the F1, precision and recall scores for each class to be predicted in the brain scans (e.g., Brodersen et al.,
Finally, StLe-minded investigators use learning curves (Abu-Mostafa et al.,
In sum, the ClSt and StLe communities rely on diagnostic metrics that are largely incongruent and may therefore not lend themselves for direct comparison in all practical analysis settings.
Case study four: out-of-sample generalization and subsequent classical inference
Vignette: The investigator is interested in potential differences in brain volume that are associated with an individual's age (continuous target variable). A LASSO (often considered as StLe arsenal) is computed on the voxel-based morphometry data from the brain's gray matter of the 1,200-subject HCP release (Human Connectome Project; Van Essen et al., 2012). This L1-penalized residual-sum-of-squares regression performs automatic variable selection (i.e., effectively eliminates coefficients by setting them to zero) on all gray-matter voxels' volume information in a high-dimensional regime (i.e., no mass-univariate analysis). Assessing generalization performance of different sparse models using 5-fold cross-validation yields the non-zero coefficients for few brain voxels whose volumetric information is most predictive of an individual's age.
Question: How can the investigator perform classical inference to know which of the gray-matter voxels selected to be predictive for biological age are statistically significant?
This is an important concern because most statistical methods currently applied to large datasets perform some explicit or implicit form of variable selection (Jenatton et al., 2011; Committee on the Analysis of Massive Data et al.,
Second, the LASSO has been introduced as an elegant solution to the combinatorial problem of what subset of gray-matter voxels is sufficient for predicting an individual's age by automatic variable selection (Tibshirani, 1996). Computing voxel-wise p-values would recast this high-dimensional pattern-learning setting (i.e., considering all brain voxels at once) into a mass-univariate hypothesis-testing problem (i.e., considering one voxel after the other) where relevance would be computed independently for each voxel and correction for multiple comparisons would become necessary. Yet, recasting into the mass-univariate setting would ignore the sophisticated selection process that led to the predictive model with a reduced number of variables (Wu et al., 2009). Put differently, variable selection via the LASSO is itself a stochastic process that is however not accounted for by the theoretical guarantees of classical inference for statistical significance (Berk et al.,
Third, the portrayed conflict between more exploratory model selection by cross-validation (StLe) and more confirmatory classical inference (ClSt) is currently at the frontier of statistical development (Loftus, 2015; Taylor and Tibshirani, 2015). New methods for so-called post-selection inference (or selective inference) allow computing p-values for a set of features that have previously been chosen to be meaningful predictors by some criterion, one example being sparsity-incuding prediction algorithms such as LASSO. According to the theory of ClSt, the statistical model is to be chosen before visiting the data. Classical statistical tests and confidence intervals therefore become invalidated and the p-values become optimistically biased (Berk et al.,
In sum, in many analysis settings, the same data should typically not be used to first apply supervised learning algorithms for automatic selection of the most predictive variables and to then test for statistical significance of the variables already found to be most predictive based on these data points. The recent developments for post-selection inference can be viewed as an attempt to reconcile certain aspects of how the StLe and ClSt paradigms draw conclusions from data.
Case study five: classical inference and subsequent out-of-sample generalization
Vignette: The investigator is interested in potential brain structure differences that are associated with an individual's gender (categorical target variable) in the voxel-based morphometry data of the 1,200-subject HCP release (Human Connectome Project; Van Essen et al., 2012). First, the >100,000 voxels per brain scan are reduced to the most important 10,000 voxels to lower the computational cost and facilitate estimation of a prediction model. To this end, ANOVA (univariate test for statistical significance belonging to ClSt) is initially used to obtain a ranking of the most relevant 10,000 features from the gray matter. This selects the 10,000 out of the original >100,000 voxel variables with highest variance explaining volume differences between males and females (i.e., the gender information associated with each brain scan is used in the univariate test). Second, support vector machine classification (“multivariate” pattern-learning algorithm belonging to StLe) is performed by cross-validation on a feature space with the 10,000 preselected gray-matter measurements to predict the gender from each subject's brain scan.
Question: Is an analysis pipeline with univariate classical inference and subsequent high-dimensional prediction valid if both steps rely on gender as the target variables?
The implications of feature engineering procedures applied before training a learning algorithm is a frequent concern and can require subtle answers (Guyon and Elisseeff, 2003; Kriegeskorte et al., 2009; Lemm et al., 2011; Hanke et al., 2015). In most applications of predictive models the large majority of brain voxels will not be very informative (Brodersen et al.,
At the core of this explanation is the goal of cross-validation to yield out-of-sample estimates. In stark contrast, remember that null-hypothesis testing yields in-sample estimates as it needs all available data points to take its decision. Using the class labels for a variable selection step just before null-hypothesis testing on a same data sample would invalidate the null hypothesis (Kriegeskorte et al., 2009, 2010). Consequently, in a ClSt regime, using class information to select variables before null-hypothesis testing will incur an instance of double-dipping (or circular analysis). This also occurs when, for instance, first correlating a behavioral measure with brain activity and then using the identified subset of brain voxels for a second correlation analysis with that same behavioral measurement (Lieberman et al., 2009; Vul et al., 2009). In this scenario, voxels are submitted to two statistical tests with the same goal in a nested, non-independent fashion (Freedman, 1983). This corrupts the validity of the null hypothesis on which the reported test results conditionally depend.
Regarding interpretation of the results, the classifier will miss some brain voxels that only carry relevant information when considered in voxel ensembles. This is because the ANOVA filter has kept voxels that are independently relevant (Brodersen et al.,
Finally, it is interesting to consider that ANOVA-mediated feature selection to a subset of p < 500 voxel variables would reduce the “wide” neuroimaging data (“n < < p” setting) down to “long” neuroimaging data with fewer features than observations (“n > p” setting) given the n = 500 subjects (Wainwright, 2014). This allows recasting the StLe regime into a ClSt regime in order to fit a GLM and perform classical statistical tests instead of training a predictive classification algorithm (Brodersen et al.,
In sum, in many analysis settings, prediction algorithms can be trained after choosing the input variables most significantly associated with an explanatory target variable if the initial classical inference (p-values) is performed only in the training set and the ensuing evaluation of algorithm generalization (prediction performance) is performed on the independent test set.
Case study six: structure discovery by clustering algorithms
Vignette: Each functionally specialized region in the human brain probably has a unique set of long-range connections (Passingham et al., 2002). This notion has prompted connectivity-based parcellation methods in neuroimaging that segregate an ROI (can be locally circumscribed or brain global; Eickhoff et al., 2015) into distinct cortical modules (Behrens et al.,
Question: Is it possible to decide whether the obtained brain clusters are statistically significant?
In essence, the aim of connectivity-guided brain parcellation is to find useful, simplified structure by imposing circumscribed compartments on brain topography (Yeo et al., 2011; Smith et al., 2013; Frackowiak and Markram, 2015). This is typically achieved by using k-means, hierarchical, Ward, or spectral clustering algorithms (Thirion et al., 2014; Eickhoff et al., 2015). Putting on the ClSt hat, an ROI clustering result would be deemed statistically significant if the obtained data are incompatible with the null hypothesis that the investigator seeks to reject (Everitt, 1979; Halkidi et al., 2001). Choosing a test statistic for clustering solutions to obtain p-values is difficult (Vogelstein et al., 2014) because of the need to find a meaningful null hypothesis to test against (Jain et al., 1999). Put differently, for classical inference based on statistical hypothesis testing one may need to pick an arbitrary null hypothesis to falsify. It follows that neither the ClSt notions of effect size and power do seem to apply in the case of brain parcellation (also a frequent question by paper reviewers). Instead of classical inference to formally test for a particular structure in the clustering results, the investigator actually needs to resort to exploratory approaches that discover and assess structure in the neuroimaging data (Tukey, 1962; Efron and Tibshirani, 1991; Hastie et al., 2001). Although statistical methods span a continuum between the two poles of ClSt and StLe, finding a clustering model with the highest fit in the sense of explaining the regional connectivity differences at hand is perhaps more naturally situated in the StLe community.
Putting on the StLe hat, the investigator realizes that the problem of brain parcellation constitutes an unsupervised learning setting without any target variable y to predict (e.g., cognitive tasks, the age or gender of the participants). The learning problem does therefore not consist in estimating a supervised predictive model y = f(X), but to estimate an unsupervised descriptive model for the connectivity data X themselves. Solving such unsupervised estimation problems is generally recognized to be ill-posed because it is generally unclear what the best way is to quantify how well relevant structure has been captured and what notion of “relevance” is most pertinent (Hastie et al., 2001; Ghahramani, 2004; Bishop,
Evidently, the discovered set of connectivity-derived clusters only represent hints to candidate brain modules. Their “existence” in neurobiology requires further scrutiny (Thirion et al., 2014; Eickhoff et al., 2015). Nevertheless, such clustering solutions provide important means to narrow down high-dimensional neuroimaging data. Preliminary clustering results broaden the space of research hypotheses that the investigator can articulate. For instance, unexpected discovery of a candidate brain region (cf. Mars et al., 2012; zu Eulenburg et al., 2012) can provide an argument for future experimental investigations. Brain parcellation can thus be viewed as an exploratory unsupervised method outlining relevant structure in neuroimaging data that can subsequently be tested as research hypotheses in targeted future neuroimaging studies on classical inference or out-of-sample generalization.
In sum, in most analysis settings, quantifying the importance of clustering solutions is inherently ill-posed because, without an explanatory target variable, many different low-dimensional reexpressions of high-dimensional input data can be useful. Choosing the right variant among the possible dimensionality reductions by clustering algorithms alone can typically not be done based on extrapolation metrics from ClSt (p-values, effect size, power) or StLe (out-of-sample prediction performance, learning curves).
Conclusion
A novel scientific fact about the brain is only valid in the context of the complexity restrictions that have been imposed on the studied phenomenon during the investigation (Box,
Statements
Author contributions
The author confirms being the sole contributor of this work and approved it for publication.
Conflict of interest
The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Footnotes
1.^“Data Science and Statistics: different worlds?” (Panel at Royal Statistical Society UK, March 2015) (https://www.youtube.com/watch?v=C1zMUjHOLr4”)
2.^“50 years of Data Science” (David Donoho, Tukey Centennial workshop, USA, September 2015)
3.^“Are ML and Statistics Complementary?” (Max Welling, 6th IMS-ISBA meeting, December 2015)
4.^In the supervised setting, there is no a priori distinction between learning algorithms evaluated by out-of-sample prediction error. In the optimization setting of finite spaces, all algorithms searching an extremum perform identically when averaged across possible cost functions. (http://www.no-free-lunch.org/)
5.^“If you torture the data enough, nature will always confess.” (Coase R. H.,
6.^“Once applied only to the selected few, the interpretation of the usual measures of uncertainty do not remain intact directly, unless properly adjusted.” (Yoav Benjamini)
References
1
Abu-MostafaY. S.Magdon-IsmailM.LinH. T. (2012). Learning from Data. AMLBook.
2
AltmanD. G.BlandJ. M. (1994). Statistics notes: diagnostic tests 2: predictive values. BMJ309:102. 10.1136/bmj.309.6947.102
3
AmuntsK.LepageC.BorgeatL.MohlbergH.DickscheidT.RousseauM. E.et al. (2013). BigBrain: an ultrahigh-resolution 3D human brain model. Science340, 1472–1475. 10.1126/science.1235381
4
AndersonD. R.BurnhamK. P.ThompsonW. L. (2000). Null hypothesis testing: problems, prevalence, and an alternative. J. Wildl. Manage.912–923. 10.2307/3803199
5
AndersonM. L. (2010). Neural reuse: a fundamental organizational principle of the brain. Behav. Brain Sci.33, 245–266; discussion 266–313. 10.1017/S0140525X10000853
6
ArbabshiraniM. R.PlisS.SuiJ.CalhounV. D. (2017). Single subject prediction of brain disorders in neuroimaging: promises and pitfalls. Neuroimage145, 137–165. 10.1016/j.neuroimage.2016.02.079
7
AverbeckB. B.LathamP. E.PougetA. (2006). Neural correlations, population coding and computation. Nat. Rev. Neurosci.7, 358–366. 10.1038/nrn1888
8
BachF. (2014). Breaking the curse of dimensionality with convex neural networks. arXiv:1412.8690.
9
BehrensT. E.Johansen-BergH.WoolrichM. W.SmithS. M.Wheeler-KingshottC. A.BoulbyP. A.et al. (2003). Non-invasive mapping of connections between human thalamus and cortex using diffusion imaging. Nat. Neurosci.6, 750–757. 10.1038/nn1075
10
BellecP.Rosa-NetoP.LytteltonO. C.BenaliH.EvansA. C. (2010). Multi-level bootstrap analysis of stable clusters in resting-state fMRI. Neuroimage51, 1126–1139. 10.1016/j.neuroimage.2010.02.082
11
BellmanR. E. (1961). Adaptive Control Processes: A Guided Tour. Princeton, NJ: Princeton University Press.
12
BengioY. (2014). Evolving culture versus local minima, in Growing Adaptive Machines, Vol. 557, eds KowaliwT.BredecheN.DoursatR. (Berlin; Heidelberg: Springer), 109–138.
13
BengioY.CourvilleA.VincentP. (2013). Representation learning: a review and new perspectives. IEEE Trans. Pattern Anal. Mach. Intell.35, 1798–1828. 10.1109/TPAMI.2013.50
14
BerkR.BrownL.BujaA.ZhangK.ZhaoL. (2013). Valid post-selection inference. Ann. Stat.41, 802–837. 10.1214/12-AOS1077
15
BerksonJ. (1938). Some difficulties of interpretation encountered in the application of the chi-square test. J. Am. Stat. Assoc.33, 526–536. 10.1080/01621459.1938.10502329
16
BishopC. M. (2006). Pattern Recognition and Machine Learning. Heidelberg: Springer.
17
BishopC. M.LasserreJ. (2007). Generative or discriminative? getting the best of both worlds. Bayesian Stat.8, 3–24.
18
BleiD. M.SmythP. (2017). Science and data science. Proc. Natl. Acad. Sci. U.S.A. 114, 8689–8692. 10.1073/pnas.1702076114
19
BoxG. E. P. (1976). Science and statistics. J. Am. Stat. Assoc.71, 791–799. 10.1080/01621459.1976.10480949
20
BreimanL. (2001). Statistical modeling: the two cultures. Stat. Sci.16, 199–231. 10.1214/ss/1009213726
21
BrodersenK. H. (2009). Decoding Mental Activity from Neuroimaging Data — the Science Behind Mind-Reading, Vol. 4. Oxford: The New Collection.
22
BrodersenK. H.DaunizeauJ.MathysC.ChumbleyJ. R.BuhmannJ. M.StephanK. E. (2013). Variational Bayesian mixed-effects inference for classification studies. Neuroimage76, 345–361. 10.1016/j.neuroimage.2013.03.008
23
BrodersenK. H.HaissF.OngC. S.JungF.TittgemeyerM.BuhmannJ. M.et al. (2011a). Model-based feature construction for multivariate decoding. Neuroimage56, 601–615. 10.1016/j.neuroimage.2010.04.036
24
BrodersenK. H.SchofieldT. M.LeffA. P.OngC. S.LomakinaE. I.BuhmannJ. M.et al. (2011b). Generative embedding for model-based classification of fMRI data. PLoS Comput. Biol.7:e1002079. 10.1371/journal.pcbi.1002079
25
BühlmannP.Van De GeerS. (2011). Statistics for High-Dimensional Data: Methods, Theory and Applications.New York, NY: Springer Science & Business Media.
26
BurnhamK. P.AndersonD. R. (2014). P values are only an index to evidence: 20th-vs. 21st-century statistical science. Ecology95, 627–630. 10.1890/13-1066.1
27
BzdokD.EickenbergM.GriselO.ThirionB.VaroquauxG. (2015). Semi-supervised factored logistic regression for high-dimensional neuroimaging Data, in NIPS'15 Proceedings of the 28th International Conference on Neural Information Processing Systems, (Cambridge, MA), 3348–3356.
28
BzdokD.EickenbergM.VaroquauxG.ThirionB. (2017). Hierarchical region-network sparsity for high-dimensional inference in brain imaging, in International Conference on Information Processing in Medical Imaging (IPMI) (Boone, NC).
29
BzdokD.VaroquauxG.GriselO.EickenbergM.PouponC.ThirionB. (2016). Formal models of the network co-occurrence underlying mental operations. PLoS Comput. Biol.12:e1004994. 10.1371/journal.pcbi.1004994
30
BzdokD.YeoB. T. T. (2017). Inference in the age of big data: future perspectives on neuroscience. Neuroimage155, 549–564. 10.1016/j.neuroimage.2017.04.061
31
CasellaG.BergerR. L. (2002). Statistical Inference. Pacific Grove, CA: Duxbury.
32
ChamberlinT. C. (1890). The method of multiple working hypotheses. Science15, 92–96.
33
ChambersJ. M. (1993). Greater or lesser statistics: a choice for future research. Stat. Comput.3, 182–184. 10.1007/BF00141776
34
ChoiY.TaylorJ.TibshiraniR. (2014). Selecting the number of principal components: estimation of the true rank of a noisy matrix. arXiv:1410.8260.
35
ChowS. L. (1998). Precis of statistical significance: rationale, validity, and utility. Behav. Brain Sci.21, 169–194; discussion 194–239. 10.1017/S0140525X98001162
36
ChristoffK.IrvingZ. C.FoxK. C. R.SprengR. N.Andrews-HannaJ. R. (2016). Mind-wandering as spontaneous thought: a dynamic framework. Nat. Rev. Neurosci.17, 718–731. 10.1038/nrn.2016.113
37
ChumbleyJ. R.FristonK. J. (2009). False discovery rate revisited: FDR and topological inference using Gaussian random fields. Neuroimage44, 62–70. 10.1016/j.neuroimage.2008.05.021
38
ClevelandW. S. (2001). Data science: an action plan for expanding the technical areas of the field of statistics. Int. Stat. Rev.69, 21–26. 10.1111/j.1751-5823.2001.tb00477.x
39
CoaseR. H. (1982). How Should Economists Choose? The G. Warren Nutter Lectures in Political Economy.Washington, DC: American Enterprise Institute for Public Policy Research.
40
CohenJ. (1977). Statistical Power Analysis for the Behavioral Sciences. Hillsdale, NJ: Lawrence Erlbaum Associates, Inc.
41
CohenJ. (1990). Things I have learned (so far). Am. Psychol.45:1304. 10.1037/0003-066X.45.12.1304
42
CohenJ. (1992). A power primer. Psychol. Bull.112:155. 10.1037/0033-2909.112.1.155
43
CohenJ. (1994). The Earth Is Round (p < 0.05). Am. Psychol.49, 997–1003. 10.1037/0003-066X.49.12.997
44
Committee on the Analysis of Massive Data, Committee on Applied and Theoretical Statistics, Board on Mathematical Sciences and their Applications, Division on Engineering and Physical Sciences, and National Research Council. (2013). Frontiers in Massive Data Analysis. Washington, DC: The National Academies Press.
45
CowlesM.DavisC. (1982). On the origins of the.05 level of statistical significance. Am. Psychol.37, 553–558.
46
CoxD. D.DeanT. (2014). Neural networks and neuroscience-inspired computer vision. Curr. Biol.24, R921–R929. 10.1016/j.cub.2014.08.026
47
CoxD. D.SavoyR. L. (2003). Functional magnetic resonance imaging (fMRI) “brain reading”: detecting and classifying distributed patterns of fMRI activity in human visual cortex. Neuroimage19, 261–270. 10.1016/S1053-8119(03)00049-1
48
CoxD. R. (1975). A note on data-splitting for the evaluation of significance levels. Biometrika62, 441–444. 10.1093/biomet/62.2.441
49
CummingG. (2009). Inference by eye: reading the overlap of independent confidence intervals. Stat. Med.28, 205–220. 10.1002/sim.3471
50
DavatzikosC. (2004). Why voxel-based morphometric analysis should be used with great caution when characterizing group differences. Neuroimage23, 17–20. 10.1016/j.neuroimage.2004.05.010
51
DavisJ.GoadrichM. (2006). The relationship between Precision-Recall and ROC curves,” in Proceedings of the 23rd International Conference on Machine Learning (Pittsburgh, PA: ACM), 233–240.
52
de BrebissonA.MontanaG. (2015). Deep neural networks for anatomical brain segmentation. arXiv:1502.02445.
53
DemšarJ. (2006). Statistical comparisons of classifiers over multiple data sets. J. Mach. Learn. Res.7, 1–30.
54
DerrfussJ.MarR. A. (2009). Lost in localization: the need for a universal coordinate database. Neuroimage48, 1–7. 10.1016/j.neuroimage.2009.01.053
55
de-WitL.AlexanderD.EkrollV.WagemansJ. (2016). Is neuroimaging measuring information in the brain?Psychon. Bull. Rev.23, 1415–1428. 10.3758/s13423-016-1002-0
56
DomingosP. (2012). A few useful things to know about machine learning. Commun. ACM55, 78–87. 10.1145/2347736.2347755
57
DonohoD. (2015). 50 years of data science, in Based on a Presentation at the Tukey Centennial Workshop (Princeton: NJ).
58
EfronB. (1979). Bootstrap methods: another look at the jackknife. Ann. Stat.7, 1–26. 10.1214/aos/1176344552
59
EfronB. (2012). Large-Scale Inference: Empirical Bayes Methods for Estimation, Testing, and Prediction.Cambridge, UK: Cambridge University Press.
60
EfronB.HastieT. (2016). Computer-Age Statistical Inference. Cambridge, UK: Cambridge University Press.
61
EfronB.TibshiraniR. J. (1991). Statistical data analysis in the computer age. Science253, 390–395. 10.1126/science.253.5018.390
62
EfronB.TibshiraniR. J. (1994). An Introduction to the Bootstrap. London, UK: CRC press.
63
EickhoffS. B.BzdokD.LairdA. R.RoskiC.CaspersS.ZillesK.et al. (2011). Co-activation patterns distinguish cortical modules, their connectivity and functional differentiation. Neuroimage57, 938–949. 10.1016/j.neuroimage.2011.05.021
64
EickhoffS. B.ThirionB.VaroquauxG.BzdokD. (2015). Connectivity-based parcellation: critique and implications. Hum. Brain Mapp.36, 4771–479210.1002/hbm.22933
65
EickhoffS.TurnerJ. A.NicholsT. E.Van HornJ. D. (2016). Sharing the wealth: neuroimaging data repositories. Neuroimage124, 1065–1068. 10.1016/j.neuroimage.2015.10.079
66
EstesW. K. (1997). On the communication of information by displays of standard errors and confidence intervals. Psychon. Bull. Rev.4, 330–341. 10.3758/BF03210790
67
EverittB. S. (1979). Unresolved problems in cluster analysis. Biometrics35, 169–181. 10.2307/2529943
68
FergusonC. J. (2009). An effect size primer: a guide for clinicians and researchers. Prof. Psychol.40:532. 10.1037/a0015808
69
FeyerabendP. (1975). Against Method: Outline of an Anarchist Theory of Knowledge. London: New Left Books.
70
FisherR. A. (1925). Statistical Methods of Research Workers. London: Oliver and Boyd.
71
FisherR. A. (1935). The Design of Experiments.Edinburgh: Oliver and Boyd.
72
FisherR. A.MackenzieW. A. (1923). Studies in crop variation. II. The manurial response of different potato varieties. J. Agric. Sci.13, 311–320. 10.1017/S0021859600003592
73
FithianW.SunD.TaylorJ. (2014). Optimal inference after model selection. arXiv:1410.2597.
74
FleckL.SchäferL.SchnelleT. (1935). Entstehung und Entwicklung einer Wissenschaftlichen Tatsache. Basel: Schwabe.
75
FoxP. T.LancasterJ. L.LairdA. R.EickhoffS. B. (2014). Meta-analysis in human neuroimaging: computational modeling of large-scale databases. Annu. Rev. Neurosci.37, 409–434. 10.1146/annurev-neuro-062012-170320
76
FrackowiakR.MarkramH. (2015). The future of human cerebral cartography: a novel approach. Philos. Trans. R. Soc. Lond. B Biol. Sci.370:20140171. 10.1098/rstb.2014.0171
77
FreedmanD. A. (1983). A note on screening regression equations. Am. Stat.37, 152–155.
78
FriedmanJ. H. (1998). Data mining and statistics: what's the connection?Comput. Sci. Stat.29, 3–9.
79
FriedmanJ. H. (2001). The role of statistics in the data revolution?Int. Stat. Rev.69, 5–10.
80
FrimanO.CedefamnJ.LundbergP.BorgaM.KnutssonH. (2001). Detection of neural activity in functional MRI using canonical correlation analysis. Magn. Reson. Med.45, 323–330. 10.1002/1522-2594(200102)45:2<323::AID-MRM1041>3.0.CO;2-#
81
FristonK. J. (2006). Statistical Parametric Mapping: The Analysis of Functional Brain Images. Amsterdam: Academic Press.
82
FristonK. J. (2009). Modalities, modes, and models in functional neuroimaging. Science326, 399–403. 10.1126/science.1174521
83
FristonK. J. (2012). Ten ironic rules for non-statistical reviewers. Neuroimage61, 1300–1310. 10.1016/j.neuroimage.2012.04.018
84
FristonK. J.ChuC.Mourao-MirandaJ.HulmeO.ReesG.PennyW.et al. (2008). Bayesian decoding of brain images. Neuroimage39, 181–205. 10.1016/j.neuroimage.2007.08.013
85
FristonK. J.HolmesA. P.WorsleyK. J.PolineJ. P.FrithC. D.FrackowiakR. S. (1994). Statistical parametric maps in functional imaging: a general linear approach. Hum. Brain Mapp.2, 189–210. 10.1002/hbm.460020402
86
FristonK. J.LiddleP. F.FrithC. D.HirschS. R.FrackowiakR. S. J. (1992). The left medial temporal region and schizophrenia. Brain115, 367–382. 10.1093/brain/115.2.367
87
FristonK. J.PriceC. J.FletcherP.MooreC.FrackowiakR. S. J.DolanR. J. (1996). The trouble with cognitive subtraction. Neuroimage4, 97–104. 10.1006/nimg.1996.0033
88
GabrieliJ. D.GhoshS. S.Whitfield-GabrieliS. (2015). Prediction as a humanitarian and pragmatic contribution from human cognitive neuroscience. Neuron85, 11–26. 10.1016/j.neuron.2014.10.047
89
GenoveseC. R.LazarN. A.NicholsT. (2002). Thresholding of statistical maps in functional neuroimaging using the false discovery rate. Neuroimage15, 870–878. 10.1006/nimg.2001.1037
90
GhahramaniZ. (2004). Unsupervised learning, in Advanced Lectures on Machine Learning, eds BousquetO.von LuxburgU.RätschG. (Berlin; Heidelberg: Springer), 72–112.
91
GhahramaniZ. (2015). Probabilistic machine learning and artificial intelligence. Nature521, 452–459. 10.1038/nature14541
92
GigerenzerG. (1993). The superego, the ego, and the id in statistical reasoning, in A Handbook for Data Analysis in the Behavioral Sciences: Methodological issues, eds KerenG.LewisC. (Hillslade, NJ: Lawrence Erlbaum Associates), 311–339.
93
GigerenzerG. (2004). Mindless statistics. J. Soc. Econ.33, 587–606. 10.1016/j.socec.2004.09.033
94
GigerenzerG.MurrayD. J. (1987). Cognition as Intuitive Statistics. Hillsdale, NJ: Erlbaum.
95
GiraudC. (2014). Introduction to High-Dimensional Statistics. London, UK: CRC Press.
96
GläscherJ.AdolphsR.DamasioH.BecharaA.RudraufD.CalamiaM.et al. (2012). Lesion mapping of cognitive control and value-based decision making in the prefrontal cortex. Proc. Natl. Acad. Sci. U.S.A.109, 14681–14686. 10.1073/pnas.1206608109
97
GollandP.FischlB. (2003). Permutation tests for classification: towards statistical significance in image-based studies, in Information Processing in Medical Imaging, eds TaylorC.NobleJ. A. (Berlin; Heidelberg: Springer), 330–341.
98
GoodfellowI. J.BengioY.CourvilleA. (2016). Deep Learning. Cambridge, MA: MIT Press.
99
GoodmanS. N. (1999). Toward evidence-based medical statistics. 1: the P value fallacy. Ann. Int. Med.130, 995–1004. 10.7326/0003-4819-130-12-199906150-00008
100
GradyC. L.HaxbyJ. V.SchapiroM. B.Gonzalez-AvilesA.KumarA.BallM. J.et al. (1990). Subgroups in dementia of the Alzheimer type identified using positron emission tomography. J. Neuropsychiatry Clin. Neurosci.2, 373–384. 10.1176/jnp.2.4.373
101
GreenwaldA. G. (2012). There is nothing so theoretical as a good method. Perspect. Psychol. Sci.7, 99–108. 10.1177/1745691611434210
102
GüçlüU.van GervenM. A. J. (2015). Deep neural networks reveal a gradient in the complexity of neural representations across the ventral stream. J. Neurosci.35, 10005–10014. 10.1523/JNEUROSCI.5023-14.2015
103
GuyonI.ElisseeffA. (2003). An introduction to variable and feature selection. J. Mach. Learn. Res.3, 1157–1182.
104
GuyonI.WestonJ.BarnhillS.VapnikV. (2002). Gene selection for cancer classification using support vector machines. Mach. Learn.46, 389–422. 10.1023/A:1012487302797
105
HalkidiM.BatistakisY.VazirgiannisM. (2001). On clustering validation techniques. J. Intell. Inf. Syst.17, 107–145. 10.1023/A:1012801612483
106
HallE. T. (1976). Beyond Culture. New York, NY: Anchor Books.
107
HandlJ.KnowlesJ.KellD. B. (2005). Computational cluster validation in post-genomic data analysis. Bioinformatics21, 3201–3212. 10.1093/bioinformatics/bti517
108
HankeM.HalchenkoY. O.OosterhofN. N. (2015). PyMVPA Manuel. Available online at: http://www.pymvpa.org/
109
HankeM.HalchenkoY. O.SederbergP. B.HansonS. J.HaxbyJ. V.PollmannS. (2009). PyMVPA: a python toolbox for multivariate pattern analysis of fMRI data. Neuroinformatics7, 37–53. 10.1007/s12021-008-9041-y
110
HansonS. J.HalchenkoY. O. (2008). Brain reading using full brain support vectormachines for object recognition: there is no “Face” Identification Area. Neural Comput.20, 486–503. 10.1162/neco.2007.09-06-340
111
HansonS. J.MatsukaT.HaxbyJ. V. (2004). Combinatorial codes in ventral temporal lobe for object recognition: Haxby (2001) revisited: is there a “face” area?Neuroimage23, 156–166. 10.1016/j.neuroimage.2004.05.020
112
HastieT.TibshiraniR.FriedmanJ. (2001). The Elements of Statistical Learning. Springer Series in Statistics. Heidelberg: Springer.
113
HastieT.TibshiraniR.WainwrightM. (2015). Statistical Learning with Sparsity. The Lasso and Generalizations.London, UK: CRC Press.
114
HaxbyJ. V. (2012). Multivariate pattern analysis of fMRI: the early beginnings. Neuroimage62, 852–855. 10.1016/j.neuroimage.2012.03.016
115
HaxbyJ. V.GobbiniM. I.FureyM. L.IshaiA.SchoutenJ. L.PietriniP. (2001). Distributed and overlapping representations of faces and objects in ventral temporal cortex. Science293, 2425–2430. 10.1126/science.1063736
116
HaynesJ.-D. (2015). A primer on pattern-based approaches to fMRI: principles, pitfalls, and perspectives. Neuron87, 257–270. 10.1016/j.neuron.2015.05.025
117
HaynesJ. D.ReesG. (2005). Predicting the orientation of invisible stimuli from acitvity in human primary visual cortex. Nat. Neurosci.8, 686–691. 10.1038/nn1445
118
HaynesJ. D.ReesG. (2006). Decoding mental states from brain activity in humans. Nat. Rev. Neurosci.7, 523–534. 10.1038/nrn1931
119
HenkeN.BughinJ.ChuiM.ManyikaJ.SalehT.WisemanB.et al. (2016). The Age of Analytics: Competing in a data-driven world. Technical Report, McKinsey Global Institute.
120
HintonG. E.SalakhutdinovR. R. (2006). Reducing the dimensionality of data with neural networks. Science313, 504–507. 10.1126/science.1127647
121
IoannidisJ. P. (2005). Why most published research findings are false. PLoS Med.2:e124. 10.1371/journal.pmed.0020124
122
JainA. K.MurtyM. N.FlynnP. J. (1999). Data clustering: a review. ACN Comput. Surv.31, 264–323. 10.1145/331499.331504
123
JamalabadiH.AlizadehS.SchönauerM.LeiboldC.GaisS. (2016). Classification based hypothesis testing in neuroscience: below-chance level classification rates and overlooked statistical properties of linear parametric classifiers. Hum. Brain Mapp.37, 1842–1855. 10.1002/hbm.23140
124
JamesG.WittenD.HastieT.TibshiraniR. (2013). An Introduction to Statistical Learning. New York, NY: Springer.
125
JenattonR.AudibertJ.-Y.BachF. (2011). Structured variable selection with sparsity-inducing norms. J. Mach. Learn. Res.12, 2777–2824.
126
JordanM. I.MitchellT. M. (2015). Machine learning: trends, perspectives, and prospects. Science349, 255–260. 10.1126/science.aaa8415
127
KamitaniY.SawahataY. (2010). Spatial smoothing hurts localization but not information: pitfalls for brain mappers. Neuroimage49, 1949–1952. 10.1016/j.neuroimage.2009.06.040
128
KamitaniY.TongF. (2005). Decoding the visual and subjective contents of the human brain. Nat. Neurosci.8, 679–685. 10.1038/nn1444
129
KandelE. R.MarkramH.MatthewsP. M.YusteR.KochC. (2013). Neuroscience thinks big (and collaboratively). Nat. Rev. Neurosci.14, 659–664. 10.1038/nrn3578
130
KelleyK.PreacherK. J. (2012). On effect size. Psychol. Methods17, 137. 10.1037/a0028086
131
KingJ. R.DehaeneS. (2014). Characterizing the dynamics of mental representations: the temporal generalization method. Trends Cogn. Sci.18, 203–210. 10.1016/j.tics.2014.01.002
132
KnopsA.ThirionB.HubbardE. M.MichelV.DehaeneS. (2009). Recruitment of an area involved in eye movements during mental arithmetic. Science324, 1583–1585. 10.1126/science.1171599
133
KriegeskorteN. (2011). Pattern-information analysis: from stimulus decoding to computational-model testing. Neuroimage56, 411–421. 10.1016/j.neuroimage.2011.01.061
134
KriegeskorteN.GoebelR.BandettiniP. (2006). Information-based functional brain mapping. Proc. Natl. Acad. Sci. U.S.A.103, 3863–3868. 10.1073/pnas.0600244103
135
KriegeskorteN.LindquistM. A.NicholsT. E.PoldrackR. A.VulE. (2010). Everything you never wanted to know about circular analysis, but were afraid to ask. J. Cereb. Blood Flow Metab.30, 1551–1557. 10.1038/jcbfm.2010.86
136
KriegeskorteN.SimmonsW. K.BellgowanP. S.BakerC. I. (2009). Circular analysis in systems neuroscience: the dangers of double dipping. Nat. Neurosci.12, 535–540. 10.1038/nn.2303
137
KurzweilR. (2005). The Singularity is Near: When Humans Transcend Biology. London, UK: Penguin.
138
LakeB. M.SalakhutdinovR.TenenbaumJ. B. (2015). Human-level concept learning through probabilistic program induction. Science350, 1332–1338. 10.1126/science.aab3050
139
LeCunY.BengioY.HintonG. (2015). Deep learning. Nature521, 436–444. 10.1038/nature14539
140
LemmS.BlankertzB.DickhausT.MullerK. R. (2011). Introduction to machine learning for brain imaging. Neuroimage56, 387–399. 10.1016/j.neuroimage.2010.11.004
141
LiebermanM. D.BerkmanE. T.WagerT. D. (2009). Correlations in social neuroscience aren't Voodoo: Commentary on Vul et al. Perspect. Psychol. Sci.4, 299–307. 10.1111/j.1745-6924.2009.01128.x
142
LoA.ChernoffH.ZhengT.LoS. H. (2015). Why significant variables aren't automatically good predictors. Proc. Natl. Acad. Sci. U.S.A.112, 13892–13897. 10.1073/pnas.1518285112
143
LoftusJ. R. (2015). Selective inference after cross-validation. arXiv:1511.08866.
144
LogothetisN. K.PaulsJ.AugathM.TrinathT.OeltermannA. (2001). Neurophysiological investigation of the basis of the fMRI signal. Nature412, 150–157. 10.1038/35084005
145
ManyikaJ.ChuiM.BrownB.BughinJ.DobbsR.RoxburghC.et al. (2011). Big Data: The Next Frontier for Innovation, Competition, and Productivity. Technical Report, McKinsey Global Institute.
146
MarkramH. (2012). The human brain project. Sci. Am.306, 50–55. 10.1038/scientificamerican0612-50
147
MarsR. B.SalletJ.SchuffelgenU.JbabdiS.ToniI.RushworthM. F. (2012). Connectivity-based subdivisions of the human right “Temporoparietal Junction Area”: evidence for different areas participating in different cortical networks. Cereb. Cortex22, 1894–1903. 10.1093/cercor/bhr268
148
MillerK. L.Alfaro-AlmagroF.BangerterN. K.ThomasD. L.YacoubE.XuJ.et al. (2016). Multimodal population brain imaging in the UK Biobank prospective epidemiological study. Nat. Neurosci.19, 1523–1536. 10.1038/nn.4393
149
MisakiM.KimY.BandettiniP. A.KriegeskorteN. (2010). Comparison of multivariate classifiers and response normalizations for pattern-information fMRI. Neuroimage53, 103–118. 10.1016/j.neuroimage.2010.05.051
150
MoellerJ. R.StrotherS. C.SidtisJ. J.RottenbergD. A. (1987). Scaled subprofile model: a statistical approach to the analysis of functional patterns in positron emission tomographic data. J. Cereb. Blood Flow Metab.7, 649–658. 10.1038/jcbfm.1987.118
151
MurM.BandettiniP. A.KriegeskorteN. (2009). Revealing representational content with pattern-information fMRI–an introductory guide. Soc. Cogn. Affect. Neurosci.4, 101–109. 10.1093/scan/nsn044
152
MurphyK. P. (2012). Machine Learning: A Probabilistic Perspective. Cambridge, UK: MIT Press.
153
NaselarisT.KayK. N.NishimotoS.GallantJ. L. (2011). Encoding and decoding in fMRI. Neuroimage56, 400–410. 10.1016/j.neuroimage.2010.07.073
154
NeymanJ.PearsonE. S. (1933). On the problem of the most efficient tests for statistical hypotheses. Philos. Trans. R. Soc. A231, 289–337. 10.1098/rsta.1933.0009
155
NicholsT. E. (2012). Multiple testing corrections, nonparametric methods, and random field theory. Neuroimage62, 811–815. 10.1016/j.neuroimage.2012.04.014
156
NicholsT. E.HayasakaS. (2003). Controlling the familywise error rate in functional neuroimaging: a comparative review. Stat. Methods Med. Res.12, 419–446. 10.1191/0962280203sm341ra
157
NicholsT. E.HolmesA. P. (2002). Nonparametric permutation tests for functional neuroimaging: a primer with examples. Hum. Brain Mapp.15, 1–25. 10.1002/hbm.1058
158
NickersonR. S. (2000). Null hypothesis significance testing: a review of an old and continuing controversy. Psychol. Methods5, 241–301. 10.1037/1082-989X.5.2.241
159
NoirhommeQ.LesenfantsD.GomezF.SodduA.SchrouffJ.GarrauxG.et al. (2014). Biased binomial assessment of cross-validated estimation of classification accuracies illustrated in diagnosis predictions. Neuroimage4, 687–694. 10.1016/j.nicl.2014.04.004
160
NormanK. A.PolynS. M.DetreG. J.HaxbyJ. V. (2006). Beyond mind-reading: multi-voxel pattern analysis of fMRI data. Trends Cogn. Sci.10, 424–430. 10.1016/j.tics.2006.07.005
161
NuzzoR. (2014). Scientific method: statistical errors. Nature506, 150–152. 10.1038/506150a
162
OakesM. (1986). Statistical Inference: A Commentary for the Social and Behavioral Sciences. New York, NY: Wiley.
163
PassinghamR. E.StephanK. E.KotterR. (2002). The anatomical basis of functional localization in the cortex. Nat. Rev. Neurosci.3, 606–616. 10.1038/nrn893
164
PedregosaF.EickenbergM.CiuciuP.ThirionB.GramfortA. (2015). Data-driven HRF estimation for encoding and decoding models. Neuroimage104, 209–220. 10.1016/j.neuroimage.2014.09.060
165
PereiraF.BotvinickM. (2011). Information mapping with pattern classifiers: a comparative study. Neuroimage56, 476–496. 10.1016/j.neuroimage.2010.05.026
166
PereiraF.MitchellT.BotvinickM. (2009). Machine learning classifiers and fMRI: a tutorial overview. Neuroimage45, 199–209. 10.1016/j.neuroimage.2008.11.007
167
PernetC. R.ChauveauN.GasparC.RousseletG. A. (2011). LIMO EEG: a toolbox for hierarchical LInear MOdeling of ElectroEncephaloGraphic data. Comput. Intell. Neurosci.2011:831409. 10.1155/2011/831409
168
PlattJ. R. (1964). Strong inference: certain systematic methods of scientific thinking may produce much more rapid progress than others. Science146, 347–353. 10.1126/science.146.3642.347
169
PlisS. M.HjelmD. R.SalakhutdinovR.AllenE. A.BockholtH. J.LongJ. D.et al. (2014). Deep learning for neuroimaging: a validation study. Front. Neurosci.8:229. 10.3389/fnins.2014.00229
170
PoldrackR. A. (2006). Can cognitive processes be inferred from neuroimaging data?Trends Cogn. Sci.10, 59–63. 10.1016/j.tics.2005.12.004
171
PoldrackR. A.BakerC. I.DurnezJ.GorgolewskiK. J.MatthewsP. M.MunafòM. R.et al. (2017). Scanning the horizon: towards transparent and reproducible neuroimaging research. Nat. Rev. Neurosci.18, 115–126. 10.1038/nrn.2016.167
172
PoldrackR. A.GorgolewskiK. J. (2014). Making big data open: data sharing in neuroimaging. Nat. Neurosci.17, 1510–1517. 10.1038/nn.3818
173
PolineJ.-B.BrettM. (2012). The general linear model and fMRI: does love last forever?Neuroimage62, 871–880. 10.1016/j.neuroimage.2012.01.133
174
PopperK. (1935/2005). Logik der Forschung, 11th Edn. Tübingen: Mohr Siebeck.
175
PowersD. M. (2011). Evaluation: from precision, recall and f-measure to ROC, informedness, markedness and correlation. J. Mach. Learn. Technol.2, 37–63
176
RosenblattF. (1958). The perceptron: a probabilistic model for information storage and organization in the brain. Psychol. Rev.65, 386. 10.1037/h0042519
177
RosnowR. L.RosenthalR. (1989). Statistical procedures and the justification of knowledge in psychological science. Am. Psychol.44:1276. 10.1037/0003-066X.44.10.1276
178
RussellS. J.NorvigP. (2002). Artificial Intelligence: A Modern Approach (International Edition). London, UK: Pearson.
179
SamuelA. L. (1959). Some studies in machine learning using the game of checkers. IBM J. Res. Dev.3, 210–229. 10.1147/rd.33.0210
180
SayginZ. M.OsherD. E.KoldewynK.ReynoldsG.GabrieliJ. D.SaxeR. R. (2012). Anatomical connectivity patterns predict face selectivity in the fusiform gyrus. Nat. Neurosci.15, 321–327. 10.1038/nn.3001
181
SchefféH. (1959). The Analysis of Variance. New York, NY: Wiley.
182
SchmidtF. L. (1996). Statistical significance testing and cumulative knowledge in psychology: implications for training of researchers. Psychol. Methods1:115. 10.1037/1082-989X.1.2.115
183
SchwartzY.ThirionB.VaroquauxG. (2013). Mapping paradigm ontologies to and from the brain, in Advances in Neural Information Processing Systems.1673–1681.
184
Shalev-ShwartzS.Ben-DavidS. (2014). Understanding Machine Learning: From Theory to Algorithms. Cambridge, UK: Cambridge University Press.
185
ShmueliG. (2010). To explain or to predict?Stat. Sci.25, 289–310. 10.1214/10-STS330
186
SladekR.RocheleauG.RungJ.DinaC.ShenL.SerreD.et al. (2007). A genome-wide association study identifies novel risk loci for type 2 diabetes. Nature445, 881–885. 10.1038/nature05616
187
SmithS. M.BeckmannC. F.AnderssonJ.AuerbachE. J.BijsterboschJ.DouaudG.et al. (2013). Resting-state fMRI in the human connectome project. Neuroimage80, 144–168. 10.1016/j.neuroimage.2013.05.039
188
SmithS. M.MatthewsP. M.JezzardP. (2001). Functional MRI: An Introduction to Methods. Oxford University Press.
189
SmithS. M.NicholsT. E. (2009). Threshold-free cluster enhancement: addressing problems of smoothing, threshold dependence and localisation in cluster inference. Neuroimage44, 83–98. 10.1016/j.neuroimage.2008.03.061
190
StarkC. E.SquireL. R. (2001). When zero is not zero: the problem of ambiguous baseline conditions in fMRI. Proc. Natl. Acad. Sci. U.S.A.98, 12760–12766. 10.1073/pnas.221462998
191
TaylorJ.LockhartR.TibshiraniR. J.TibshiraniR. (2014). Exact post-selection inference for forward stepwise and least angle regression. arXiv:1401.3889.
192
TaylorJ.TibshiraniR. J. (2015). Statistical learning and selective inference. Proc. Natl. Acad. Sci. U.S.A.112, 7629–7634. 10.1073/pnas.1507583112
193
TenenbaumJ. B.KempC.GriffithsT. L.GoodmanN. D. (2011). How to grow a mind: statistics, structure, and abstraction. Science331, 1279–1285. 10.1126/science.1192788
194
ThirionB.VaroquauxG.DohmatobE.PolineJ. B. (2014). Which fMRI clustering gives good brain parcellations?Front. Neurosci.8:167. 10.3389/fnins.2014.00167
195
TibshiraniR. (1996). Regression shrinkage and selection via the lasso. J. R. Stat. Soc. B73, 267–288. 10.1111/j.1467-9868.2011.00771.x
196
TukeyJ. W. (1962). The future of data analysis. Ann. Stat.33, 1–67. 10.1214/aoms/1177704711
197
UK House of Common S.a.T (2016). The Big Data Dilemma. Committee on Applied and Theoretical Statistics.
198
VanderplasJ. (2013). The Big Data Brain Drain: Why Science is in Trouble. Pythonic Perambulations.
199
Van EssenD. C.UgurbilK.AuerbachE.BarchD.BehrensT. E.BucholzR.et al. (2012). The human connectome project: a data acquisition perspective. Neuroimage62, 2222–2231. 10.1016/j.neuroimage.2012.02.018
200
Van HornJ. D.TogaA. W. (2014). Human neuroimaging as a “Big Data” science. Brain Imaging Behav.8, 323–331. 10.1007/s11682-013-9255-y
201
VapnikV. N. (1989). Statistical Learning Theory. New York, NY: Wiley-Interscience.
202
VapnikV. N. (1996). The Nature of Statistical learnIng Theory. New York, NY: Springer.
203
VapnikV. N.KotzS. (1982). Estimation of Dependences Based on Empirical Data. New York, NY: Springer-Verlag New York.
204
VaroquauxG.ThirionB. (2014). How machine learning is shaping cognitive neuroimaging. Gigascience3:28. 10.1186/2047-217X-3-28
205
VogelsteinJ. T.ParkY.OhyamaT.KerrR. A.TrumanJ. W.PriebeC. E.et al. (2014). Discovery of brainwide neural-behavioral maps via multiscale unsupervised structure learning. Science344, 386–392. 10.1126/science.1250298
206
VulE.HarrisC.WinkielmanP.PashlerH. (2009). Puzzlingly high correlations in fMRI studies of emotion, personality, and social cognition. Perspect. Psychol. Sci.4, 274–290. 10.1111/j.1745-6924.2009.01125.x
207
WainwrightM. J. (2014). Structured regularizers for high-dimensional problems: statistical and computational issues. Annu. Rev. Stat. Appl.1, 233–253. 10.1146/annurev-statistics-022513-115643
208
WassermanL.RoederK. (2009). High dimensional variable selection. Ann. Stat.37:2178. 10.1214/08-AOS646
209
WassersteinR. L.LazarN. A. (2016). The ASA's statement on p-values: context, process, and purpose. Am. Stat.70, 129–133. 10.1080/00031305.2016.1154108
210
WolpertD. (1996). The lack of a priori distinctions between learning algorithms. Neural Comput.8, 1341–1390. 10.1162/neco.1996.8.7.1341
211
WorsleyK. J.EvansA. C.MarrettS.NeelinP. (1992). A three-dimensional statistical analysis for CBF activation studies in human brain. J. Cereb. Blood Flow Metab.12, 900–918. 10.1038/jcbfm.1992.127
212
WorsleyK. J.PolineJ.-B.FristonK. J.EvansA. C. (1997). Characterizing the response of PET and fMRI data using multivariate linear models. Neuroimage6, 305–319. 10.1006/nimg.1997.0294
213
WuT. T.ChenY. F.HastieT.SobelE.LangeK. (2009). Genome-wide association analysis by lasso penalized logistic regression. Bioinformatics25, 714–721. 10.1093/bioinformatics/btp041
214
YaminsD. L.DiCarloJ. J. (2016). Using goal-driven deep learning models to understand sensory cortex. Nat. Neurosci.19, 356–365. 10.1038/nn.4244
215
YarkoniT.BraverT. S. (2010). Cognitive neuroscience approaches to individual differences in working memory and executive control: conceptual and methodological issues, in Handbook of Individual Differences in Cognition, eds GruszkaA.MatthewsG.SzymuraB. (New York, NY: Springer), 87–107.
216
YarkoniT.PoldrackR. A.NicholsT. E.Van EssenD. C.WagerT. D. (2011). Large-scale automated synthesis of human functional neuroimaging data. Nat. Methods8, 665–670. 10.1038/nmeth.1635
217
YarkoniT.WestfallJ. (2017). Choosing prediction over explanation in psychology: lessons from machine learning. Perspect. Psychol. Sci. [Epub ahead of print]. 10.1177/1745691617693393
218
YeoB. T.KrienenF. M.CheeM. W.BucknerR. L. (2014). Estimates of segregation and overlap of functional connectivity networks in the human cerebral cortex. Neuroimage88, 212–227. 10.1016/j.neuroimage.2013.10.046
219
YeoB. T.KrienenF. M.SepulcreJ.SabuncuM. R.LashkariD.HollinsheadM.et al. (2011). The organization of the human cerebral cortex estimated by intrinsic functional connectivity. J. Neurophysiol.106, 1125–1165. 10.1152/jn.00338.2011
220
YusteR. (2015). From the neuron doctrine to neural networks. Nature Reviews Neuroscience16, 487–497. 10.1038/nrn3962
221
ZouH.HastieT. (2005). Regularization and variable selection via the elastic net. J. R. Stat. Soc.67, 301–320. 10.1111/j.1467-9868.2005.00503.x
222
zu EulenburgP.CaspersS.RoskiC.EickhoffS. B. (2012). Meta-analytical definition and functional connectivity of the human vestibular cortex. Neuroimage60, 162–16910.1016/j.neuroimage.2011.12.032
Summary
Keywords
neuroimaging, data science, epistemology, statistical inference, machine learning, p-value, Rosetta Stone
Citation
Bzdok D (2017) Classical Statistics and Statistical Learning in Imaging Neuroscience. Front. Neurosci. 11:543. doi: 10.3389/fnins.2017.00543
Received
12 April 2017
Accepted
19 September 2017
Published
06 October 2017
Volume
11 - 2017
Edited by
Yaroslav O. Halchenko, Dartmouth College, United States
Reviewed by
Matthew Brett, University of Cambridge, United Kingdom; Jean-Baptiste Poline, University of California, Berkeley, United States
Updates

Check for updates
Copyright
© 2017 Bzdok.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Danilo Bzdok danilo.bzdok@rwth-aachen.de
This article was submitted to Brain Imaging Methods, a section of the journal Frontiers in Neuroscience
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.