GENERAL COMMENTARY article

Front. Psychol., 22 November 2013

Sec. Quantitative Psychology and Measurement

Volume 4 - 2013 | https://doi.org/10.3389/fpsyg.2013.00876

Conjoint measurement of disorder prevalence, test sensitivity, and test specificity: notes on Botella, Huang, and Suero's multinomial model

  • EE

    Edgar Erdfelder *

  • MM

    Morten Moshagen

  • Department of Psychology, University of Mannheim Mannheim, Germany

Botella et al. () proposed two useful multinomial models for conjoint measurement of disorder prevalence rates in different populations (e.g., prevalence rates of dementia) and both the sensitivity and the specificity of the test used to assess this disorder (e.g., the Mini Mental State Examination, MMSE; Folstein et al., ). Their first model requires a perfect indicator of the disorder (i.e., a gold standard, GS), whereas the second model provides for indicators not perfectly correlated with the disorder (i.e., imperfect references, IR). In line with Lazarsfeld's () latent-class model, the only requirement of the latter model is local stochastic independence of the IR and the test-based classification, that is, stochastic independence of the IR and the test result within subpopulations of individuals with vs. without the disorder.

The present comment addresses two shortcomings of the IR model and suggests ways to overcome them: (1) Lack of global identifiability in general and (2) lack of local identifiability when prevalence rates are homogenous across populations.

Problem (1). As acknowledged by Botella et al. (), the IR model is not globally identifiable. There are always two sets of sensitivity and specificity parameters for both the reference (SeR and SpR, respectively) and the test (SeT and SpT, respectively) that predict exactly the same outcome probabilities and therefore cannot be distinguished on grounds of model fit [see Botella et al. (), Table 1]. Despite the lack of uniqueness in parameter estimates, Botella et al. () recommended use of the unconstrained IR model and to choose the set of parameter estimates that appears more plausible. However, besides introducing an unnecessary degree of subjectivity, a model that is consistent with parameter values incongruent with common sense is obviously too flexible and overly complex. For example, Botella et al.'s IR model allows for references and tests that are negatively correlated with the disorder under investigation, that is, for tools that measure the opposite of what they are supposed to measure. This is clearly not reasonable. In addition, their model lacks unique validity measures for both the reference and the test.

A simple way to remedy these problems is to constrain the sensitivity and specificity parameters in accordance with the two-high threshold model of detection (e.g., Snodgrass and Corwin, ; Waubert de Puiseau et al., ). In this refined model, the parameters of the IR model are reparameterized as follows:

The new parameters, DR and BR, denote validity and bias measures, respectively, for the IR [both in (0, 1)]. DR is the probability that the IR detects the true status (disorder present vs. absent), and BR represents the disorder-present bias (i.e., the probability of a positive diagnosis) given failure to detect the true status. Accordingly, the sensitivity and specificity parameter estimates of the test, SeT and SpT, are reparameterized as functions of test validity and bias parameters DT and BT, respectively. Importantly, these reparameterizations jointly imply the order constraints SeR ≥ (1 − SpR) and SeT ≥ (1 − SpT)1 so that a positive diagnosis cannot be less likely given presence than given absence of the disorder. In other words, whereas the dimensionality of the parameter space remains unchanged (as the Se and Sp parameters are replaced by D and B parameters), the refined model restricts the admissible data space. As a consequence, in contrast to Botella et al.'s IR model, the refined model excludes negative correlations of the disorder with both the IR and the test. Moreover, introducing these order constraints renders the model globally identifiable (subject to the auxiliary condition of unequal prevalence rates, see below), thereby removing any ambiguity in interpretation.

As summarized in Table 1, fitting the refined model to the data sets analyzed by Botella et al. (, Table 2) results in the same goodness-of-fit statistics as observed for the original IR model2. This shows that the order constraints are perfectly in line with the data3. However, as a consequence of exclusion of negative correlations, model flexibility as measured by cFIA is reduced for the refined model, resulting in better Minimum Description Length (MDL) indices of model fit than observed for the original IR model. An additional advantage of the refined model is that it provides unique validity and bias measures for both the reference and the test. For the MMSE data, for example, the test validity (0.736) is almost as large as the validity of the reference (0.876), although the difference in validities is statistically significant [ΔG2(1) = 8.60, p = 0.003]. Most importantly, unlike the original IR model, the refined model is globally identifiable so that there is only a single set of validity and bias estimates (and the corresponding sensitivity and specificity estimates) for both measurement tools involved (see Table 1).

Table 1

Statistic/EstimateAUDIT dataMMSE data
Original modelRefined modelOriginal modelRefined model
SeR0.996/0.000(0.996)0.876/0.000(0.876)
SpR1.000/0.004(1.000)1.000/0.124(1.000)
DR1.0000.876
BR0.000
SeT0.637/0.040(0.637)0.864/0.128(0.864)
SpT0.960/0.363(0.960)0.872/0.136(0.872)
DT0.6000.736
BT0.0980.486
G2(4)13.9913.9912.1412.14
cFIA20.118.723.021.6
MDL577.0575.61493.61492.2

Maximum likelihood parameter estimates, goodness-of-fit (G2), cFIA, and Minimum Description Length (MDL) measures for the original and the refined IR model applied to the AUDIT and the MMSE data of Botella et al. (, Table 2).

Parameter estimates in parentheses are derived from the corresponding validity and bias estimates using Equations (1) and (2). The two estimates for the original model correspond to the two maxima of the likelihood function. Note that the BR parameter for the AUDIT data is not identifiable because DR approaches the boundary of the parameter space.

Problem (2). To apply their models in situations where classification data are available from a single large study only, Botella et al. () suggested a random split of this sample in k segments and to treat these segments as if they were drawn from k different populations. However, apart from sampling error, random splits necessarily result in the same prevalence rate in each of the random segments so that the same population classification matrix must hold for each data set. In effect, there are only 3 instead of 3k independent category probabilities available, implying that both the standard and the refined IR model (with k + 4 parameters each) cannot be identifiable. Hence, random splits of a large sample will be of no help. A possible remedy is to split the sample based on a third variable that has been observed in addition to the IR and the test result (say, gender, age group, profession, or religion), provided the assumption can be made that the prevalence rates, but not the sensitivity and specificity of the test and the reference, differ between the corresponding subpopulations. Unequal prevalences in at least two subpopulations suffice to ensure local identifiability. Thus, systematic splits of a single large sample may remedy the identifiability problem whereas random splits will not.

Statements

Acknowledgments

Manuscript preparation was supported in part by grants from the Deutsche Forschungsgemeinschaft (Er 224/2-2) and the Baden-Württemberg foundation.

Footnotes

1.^This follows from (1 − SpR) = (1 − DR) · BR as implied by Equation (2) and, correspondingly, (1 − SpT) = (1 − DT) · BT.

2.^The model specification and data files used to derive the results of Table 1 with multiTree (Moshagen, ) can be requested from the first author.

3.^Although not required for the present data, it is also possible to conduct a formal ΔG difference test of the refined model against the original IR model. However, because both models include the same number of parameters and differ by a parametric order constraint only, the asymptotic distribution under the null hypothesis is a mixture of χ distributions (Iverson, ) rather than a standard χ distribution. The parametric bootstrap as implemented in multiTree (Moshagen, ) can be used to approximate this distribution.

References

Summary

Keywords

multinomial modeling, validity, diagnostic accuracy, gold standard, imperfect reference

Citation

Erdfelder E and Moshagen M (2013) Conjoint measurement of disorder prevalence, test sensitivity, and test specificity: notes on Botella, Huang, and Suero's multinomial model. Front. Psychol. 4:876. doi: 10.3389/fpsyg.2013.00876

Received

24 September 2013

Accepted

04 November 2013

Published

22 November 2013

Volume

4 - 2013

Edited by

Michel Regenwetter, University of Illinois at Urbana-Champiagn, USA

Reviewed by

Juan Botella, Universidad Autónoma de Madrid, Spain; Xiangen Hu, The University of Memphis, USA

Copyright

*Correspondence:

This article was submitted to Quantitative Psychology and Measurement, a section of the journal Frontiers in Psychology.

Disclaimer

All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.

Outline

Cite article

Copy to clipboard


Export citation file


Share article

Article metrics