Following the publication of a recent meta-analysis by Cipriani et al. (), various opinion leaders and news reports claimed that the effectiveness of antidepressants has been definitely proven (). E.g., Dr. Pariante, spokesperson for the Royal College of Psychiatrists, stated that this study “finally puts to bed the controversy on antidepressants, clearly showing that these drugs do work in lifting mood and helping most people with depression” (https://www.theguardian.com/science/2018/feb/21/the-drugs-do-work-antidepressants-are-effective-study-shows). We surely would embrace drug treatments that effectively help most people with depression, but based on work that has contested the validity of mostly industry-sponsored antidepressant trials (–) we remain skeptical about antidepressants' clinical benefits. The most recent meta-analysis indeed concludes that antidepressants are more effective than placebo but also acknowledges that risk of bias was substantial and that the mean effect size of d = 0.3 was modest (). Unfortunately, no clarification is given what this effect size means and whether it can be expected to be clinically significant in real-world routine practice. In this opinion paper we therefore ponder over how the reported effect size of d = 0.3 relates to clinical significance and how method bias undermines its validity, in order that the public, clinicians, and patients can judge for themselves whether antidepressants clearly work in most people with depression.
Statistical vs. clinical significance
Based on statistically significant drug-placebo differences, authors commonly conclude that antidepressants are effective regardless of the clinical significance of effect sizes. Cipriani et al. () even complained that there was “an undue focus on the binary and polarizing question of clinical significance” (p. 462). However, statisticians repeatedly cautioned that statistical significance does not imply practical relevance (–). A statistically significant result neither proves that the null hypothesis is false nor that the alternative hypothesis is true (, , ). Interpreting a statistically significant drug-placebo difference as evidence that drugs work is therefore a logical fallacy (). The null hypothesis is always false, as a true null-association between natural variables (i.e., d = 0.0) is nearly impossible due to residual confounding and correlational noise (, ). The American Statistical Association () formally states that “A p-value, or statistical significance, does not measure the size of an effect or the importance of a result” and they further emphasize that “Any effect, no matter how tiny, can produce a small p-value if the sample size or measurement precision is high enough …” (p. 132). With a total sample size of n = 116,477 as in the most recent meta-analysis (), it is therefore not surprising that any given drug-placebo difference, however small it may be, reaches statistical significance. Thus, since statistical significance does not imply clinical significance (, , ), readers need to consider what the reported mean effect of d = 0.3 practically means.
As shown in Figure 1, this effect size corresponds to approximately 2 points on the Hamilton Rating-Scale for Depression 17-item version (HAMD-17; range 0–52 points), but per convention a difference < 3 points or an effect size d < 0.5 (corresponding to < 4 HAMD-17 points) are considered clinically irrelevant (, ). Research suggests that drug-placebo differences < 3 points are undetectable by clinicians and that at least 7 HAMD-17 points are necessary for a clinician to detect a minimal improvement in a patient's clinical presentation (). As a result, the average treatment effect of d = 0.3 must be considered undetectable and therefore clinically insignificant in real-world routine practice. Interestingly, a previous meta-analysis by Jakobsen et al. () found comparable effect sizes, but the authors defined clinical significance a-priori and therefore questioned the real-world benefits of antidepressants. The effect sizes reported by Cipriani et al. () and Jakobsen et al. () are plotted in Figure 1.
Figure 1
Here we report Cohen's d effect sizes for the sake of completeness and because they are often reported in meta-analyses. However, we emphasize that cut-offs such as d = 0.2 (“small” effect size) or d = 0.5 (“medium” effect size) are arbitrary and should be interpreted with caution (
When based on approximately normally-distributed interval scales, d = 0.3 indicates that, first, the outcome of antidepressants and placebo overlap by 88%, second, that only 62% of participants in the antidepressant group score above the mean of the placebo group and, conversely, 38% score below the mean (referred to as Cohen's U3), and, third, that if you pick a person at random from the antidepressant group, he/she will have a minor chance of 58% to have the better outcome than a person picked at random from the placebo group (probability of 50% indicates no benefit at all) (
Addressing common objections
A frequently cited paper by Leucht et al. (
Another unsubstantiated objection is that the efficacy of antidepressants is poor due to inadequate psychometric properties of the HAMD-17 [e.g., its poor content validity (
A third objection is that critics of antidepressants unjustifiably promote psychotherapy although talk therapy is no better than pharmacotherapy. In response to these concerns we would like to state that we have also written about the limitations and biases in psychotherapy research (
The efficacy of antidepressants is overestimated
The average treatment effect detailed above, albeit minor, yet is most likely an overestimation due to various systematic biases that inflate the apparent efficacy of antidepressants, including, in particular, unblinding of outcome assessors (
First, it has consistently been shown that treatment effects are larger when the outcome is rated by unblinded assessors, thus efficacy estimates are inflated due to assessors' treatment expectancies (
Given that the mean drug-placebo difference is only about 2 HAMD-17 points, even a minor bias in symptom-ratings could fully account for antidepressants' treatment effect. Indeed, taking the observer bias into account, Gotzsche (
Conclusions
Contrary to the predominant interpretation we contend that antidepressants do not work in most patients, given that only 1 of 9 people benefit, whereas the remaining 8 are unnecessarily put at risk of adverse drug effects. To be clear, antidepressants can have strong mental and physical effects in some patients that may be considered helpful for some time (
Statements
Author contributions
MPH drafted the manuscript. MP contributed significantly in writing and critical revision.
Conflict of interest
The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
References
1.
CiprianiAFurukawaTASalantiGChaimaniAAtkinsonLZOgawaYet al. Comparative efficacy and acceptability of 21 antidepressant drugs for the acute treatment of adults with major depressive disorder: a systematic review and network meta-analysis. Lancet (2018) 391:1357–66. 10.1016/S0140-6736(17)32802-7
2.
AdlingtonK. Pop a million happy pills? Antidepressants, nuance, and the media. BMJ (2018) 360:k1069. 10.1136/bmj.k1069
3.
HengartnerMP. Methodological flaws, conflicts of interest, and scientific fallacies: Implications for the evaluation of antidepressants' efficacy and harm. Front Psychiatry (2017) 8:275. 10.3389/fpsyt.2017.00275
4.
EbrahimSBanceSAthaleAMalachowskiCIoannidisJP. Meta-analyses with industry involvement are massively published and report no caveats for antidepressants. J Clin Epidemiol. (2016) 70:155–63. 10.1016/j.jclinepi.2015.08.021
5.
MelanderHAhlqvist-RastadJMeijerGBeermannB. Evidence b(i)ased medicine–selective reporting from studies sponsored by pharmaceutical industry: review of studies in new drug applications. BMJ (2003) 326:1171–3. 10.1136/bmj.326.7400.1171
6.
TurnerEHMatthewsAMLinardatosETellRARosenthalR. Selective publication of antidepressant trials and its influence on apparent efficacy. N Engl J Med. (2008) 358:252–60. 10.1056/NEJMsa065779
7.
CiprianiASalantiGFurukawaTAEggerMLeuchtSRuheHGet al. Antidepressants might work for people with major depression: where do we go from here?Lancet Psychiat. (2018) 5:461–3. 10.1016/S2215-0366(18)30133-0
8.
CohenJ. The earth is round (P<0.05). Am Psychol (1994) 49:997–1003. 10.1037/0003-066x.50.12.1103
9.
SzucsDIoannidisJPA. When null hypothesis significance testing is unsuitable for research: a reassessment. Front Hum Neurosci. (2017) 11:390. 10.3389/fnhum.2017.00390
10.
WassersteinRLLazarNA. The ASA's statement on p-values: context, process, and purpose. Am Stat. (2016) 70:129–33. 10.1080/00031305.2016.1154108
11.
WagenmakersEJ. A practical solution to the pervasive problems of p values. Psychon B Rev. (2007) 14:779–804. 10.3758/Bf03194105
12.
HengartnerMP. What is the threshold for a clinical minimally important drug effect?BMJ Evid Based Med. (2018). 10.1136/bmjebm-2018-111056
13.
KirkRE. Practical significance: a concept whose time has come. Educ Psychol Meas. (1996) 56:746–59.
14.
JakobsenJCKatakamKKSchouAHellmuthSGStallknechtSELeth-MollerKet al. Selective serotonin reuptake inhibitors versus placebo in patients with major depressive disorder. A systematic review with meta-analysis and Trial Sequential Analysis. BMC Psychiatry (2017) 17:58. 10.1186/s12888-016-1173-2
15.
MoncrieffJKirschI. Efficacy of antidepressants in adults. BMJ (2005) 331:155–7. 10.1136/bmj.331.7509.155
16.
MoncrieffJKirschI. Empirically derived criteria cast doubt on the clinical significance of antidepressant-placebo differences. Contemp Clin Trials (2015) 43:60–2. 10.1016/j.cct.2015.05.005
17.
DurlakJA. How to select, calculate, and interpret effect sizes. J Pediatr Psychol. (2009) 34:917–28. 10.1093/jpepsy/jsp004
18.
FurukawaTACiprianiAAtkinsonLZLeuchtSOgawaYTakeshimaNet al. Placebo response rates in antidepressant trials: a systematic review of published and unpublished double-blind randomised controlled studies. Lancet Psychiat. (2016) 3:1059–66. 10.1016/S2215-0366(16)30307-8
19.
FurukawaTALeuchtS. How to obtain NNT from Cohen's d: comparison of two methods. PLoS ONE (2011) 6:e19070. 10.1371/journal.pone.0019070
20.
McCormackJKorownykC. Effectiveness of antidepressants. BMJ (2018) 360:k1073. 10.1136/bmj.k1073
21.
CarvalhoAFSharmaMSBrunoniARVietaEFavaGA. The safety, tolerability and risks associated with the use of newer generation antidepressant drugs: a critical review of the literature. Psychother Psychosom. (2016) 85:270–88. 10.1159/000447034
22.
FavaGAGattiABelaiseCGuidiJOffidaniE. Withdrawal symptoms after selective serotonin reuptake inhibitor discontinuation: a systematic review. Psychother Psychosom. (2015) 84:72–81. 10.1159/000370338
23.
FavaGABenasiGLucenteMOffidaniECosciFGuidiJ. Withdrawal symptoms after serotonin-noradrenaline reuptake inhibitor discontinuation: systematic review. Psychother Psychosom. (2018) 87:195–203. 10.1159/000491524
24.
ThaseMELarsenKGKennedySH. Assessing the ‘true' effect of active antidepressant therapy v. placebo in major depressive disorder: use of a mixture model. Br J Psychiatry (2011) 199:501–7. 10.1192/bjp.bp.111.093336
25.
LeuchtSHierlSKisslingWDoldMDavisJM. Putting the efficacy of psychiatric and general medicine medication into perspective: review of meta-analyses. Br J Psychiatry (2012) 200:97–106. 10.1192/bjp.bp.111.096594
26.
BaldessariniRJLauWKSimJSumMYSimK. Suicidal risks in reports of long-term controlled trials of antidepressants for major depressive disorder II. Int J Neuropsychopharmacol. (2017) 20:281–4. 10.1093/ijnp/pyw092
27.
BraunCBschorTFranklinJBaethgeC. Suicides and suicide attempts during long-term treatment with antidepressants: a meta-analysis of 29 placebo-controlled studies including 6,934 patients with major depressive disorder. Psychother Psychosom. (2016) 85:171–9. 10.1159/000442293
28.
HealyDWhitakerC. Antidepressants and suicide: risk-benefit conundrums. J Psychiatry Neurosci. (2003) 28:331–7.
29.
FergussonDDoucetteSGlassKCShapiroSHealyDHebertPHuttonB. BMJ (2005) 330:396. 10.1136/bmj.330.7488.396
30.
StoneMLaughrenTJonesMLLevensonMHollandPCHughesAet al. Risk of suicidality in clinical trials of antidepressants in adults: analysis of proprietary data submitted to US Food and Drug Administration. BMJ (2009) 339:b2880. 10.1136/bmj.b2880
31.
BagbyRMRyderAGSchullerDRMarshallMB. The Hamilton Depression Rating Scale: has the gold standard become a lead weight?Am J Psychiatry (2004) 161:2163–77. 10.1176/appi.ajp.161.12.2163
32.
GreenbergRPBornsteinRFGreenbergMDFisherS. A meta-analysis of antidepressant outcome under “blinder” conditions. J Consult Clin Psychol. (1992) 60:664–9.
33.
SpielmansGIGerwigK. The efficacy of antidepressants on overall well-being and self-reported depression symptom severity in youth: a meta-analysis. Psychother Psychosom. (2014) 83:158–64. 10.1159/000356191
34.
HengartnerMP. Raising awareness for the replication crisis in clinical psychology by focusing on inconsistencies in psychotherapy research: how much can we rely on published findings from efficacy trials?Front Psychol (2018) 9:256. 10.3389/fpsyg.2018.00256
35.
CuijpersPSijbrandijMKooleSLAnderssonGBeekmanATReynoldsCFIII. The efficacy of psychotherapy and pharmacotherapy in treating depressive and anxiety disorders: a meta-analysis of direct comparisons. World Psychiatry (2013) 12:137–48. 10.1002/wps.20038.
36.
CuijpersPCristeaIA. What if a placebo effect explained all the activity of depression treatments?World Psychiatry (2015) 14:310–1. 10.1002/wps.20249
37.
Biesheuvel-LeliefeldKEKokGDBocktingCLCuijpersPHollonSDvan MarwijkHWet al. Effectiveness of psychological interventions in preventing recurrence of depressive disorder: meta-analysis and meta-regression. J Affect Disord. (2015) 174:400–10. 10.1016/j.jad.2014.12.016
38.
De MaatSDekkerJSchoeversRDe JongheF. Relative efficacy of psychotherapy and pharmacotherapy in the treatment of depression: a meta-analysis. Psychother Res. (2006) 16:566–78. 10.1080/10503300600756402
39.
SpielmansGIBermanMIUsitaloAN. Psychotherapy versus second-generation antidepressants in the treatment of depression: a meta-analysis. J Nerv Ment Dis. (2011) 199:142–9. 10.1097/NMD.0b013e31820caefb
40.
MoncrieffJ. Are antidepressants as effective as claimed? No, they are not effective at all. Can J Psychiatry (2007) 52:96–7. 10.1177/070674370705200204
41.
EvenCSiobud-DorocantEDardennesRM. Critical approach to antidepressant trials. Blindness protection is necessary, feasible and measurable. Br J Psychiatry (2000) 177:47–51.
42.
HrobjartssonAThomsenASEmanuelssonFTendalBHildenJBoutronIet al. Observer bias in randomised clinical trials with binary outcomes: systematic review of trials with both blinded and non-blinded outcome assessors. BMJ (2012) 344:e1119. 10.1136/bmj.e1119
43.
KhanAFaucettJLichtenbergPKirschIBrownWA. A systematic review of comparative efficacy of treatments and controls for depression. PLoS ONE (2012) 7:e41778. 10.1371/journal.pone.0041778
44.
HrobjartssonAThomsenASEmanuelssonFTendalBHildenJBoutronIet al. Observer bias in randomized clinical trials with measurement scale outcomes: a systematic review of trials with both blinded and nonblinded assessors. CMAJ (2013) 185:E201–11. 10.1503/cmaj.120744
45.
MoncrieffJWesselySHardyR. Active placebos versus antidepressants for depression. Cochrane Database Syst Rev. (2004) 1:CD003012. 10.1002/14651858.CD003012.pub2.
46.
BarbuiCFurukawaTACiprianiA. Effectiveness of paroxetine in the treatment of acute major depression in adults: a systematic re-examination of published and unpublished data from randomized trials. CMAJ (2008) 178:296–305. 10.1503/cmaj.070693
47.
ArrollBElleyCRFishmanTGoodyear-SmithFAKenealyTBlashkiGet al. Antidepressants versus placebo for depression in primary care. Cochrane Database Syst Rev. (2009) 3:CD007954. 10.1002/14651858.CD007954.
48.
GotzschePC. Why I think antidepressants cause more harm than good. Lancet Psychiat. (2014) 1:104–6. 10.1016/S2215-0366(14)70280-9
49.
WangSMHanCLeeSJJunTYPatkarAAMasandPSet al. Efficacy of antidepressants: bias in randomized clinical trials and related issues. Expert Rev Clin Pharmacol. (2018) 11:15–25. 10.1080/17512433.2017.1377070
50.
AntonuccioDODantonWGDeNelskyGYGreenbergRPGordonJS. Raising questions about antidepressants. Psychother Psychosom. (1999) 68:3–14. 10.1159/000012304.
51.
MoncrieffJCohenD. Do antidepressants cure or create abnormal brain states?PLoS Med (2006) 3:e240. 10.1371/journal.pmed.0030240
52.
American Psychiatric Association. Diagnostic and Statistical Manual of Mental Disorders DSM-5. Washington, DC: American Psychiatric Association (2013).
53.
BregginPR. Suicidality, violence and mania caused by selective serotonin reuptake inhibitors (SSRIs): a review and analysis. Int J Risk Saf Med. (2004) 16:31–49.
54.
AndrewsPWThomsonJAJrAmstadterANealeMC. Primum non nocere: an evolutionary analysis of whether antidepressants do more harm than good. Front Psychol. (2012) 3:117. 10.3389/fpsyg.2012.00117
55.
MoretCIsaacMBrileyM. Problems associated with long-term treatment with selective serotonin reuptake inhibitors. J Psychopharmacol. (2009) 23:967–74. 10.1177/0269881108093582
56.
RichardsonKFoxCMaidmentISteelNLokeYKArthurAet al. Anticholinergic drugs and risk of dementia: case-control study. BMJ (2018) 361:k1315. 10.1136/bmj.k1315
57.
SmollerJWAllisonMCochraneBBCurbJDPerlisRHRobinsonJGet al. Antidepressant use and risk of incident cardiovascular morbidity and mortality among postmenopausal women in the Women's Health Initiative study. Arch Int Med. (2009) 169:2128–39. 10.1001/archinternmed.2009.436
58.
GafoorRBoothHPGullifordMC. Antidepressant utilisation and incidence of weight gain during 10 years' follow-up: population based cohort study. BMJ (2018) 361:k1951. 10.1136/bmj.k1951.
59.
CouplandCDhimanPMorrissRArthurABartonGHippisley-CoxJ. Antidepressant use and risk of adverse outcomes in older people: population based cohort study. BMJ (2011) 343:d4551. 10.1136/bmj.d4551
60.
MaslejMMBolkerBMRussellMJEatonKDuriskoZHollonSDet al. The mortality and myocardial effects of antidepressants are moderated by preexisting cardiovascular disease: a meta-analysis. Psychother Psychosom. (2017). 86:268–82. 10.1159/000477940
Summary
Keywords
antidepressant, meta-analysis, efficacy, effectiveness, effect size, clinical significance, method bias
Citation
Hengartner MP and Plöderl M (2018) Statistically Significant Antidepressant-Placebo Differences on Subjective Symptom-Rating Scales Do Not Prove That the Drugs Work: Effect Size and Method Bias Matter!. Front. Psychiatry 9:517. doi: 10.3389/fpsyt.2018.00517
Received
13 March 2018
Accepted
01 October 2018
Published
17 October 2018
Volume
9 - 2018
Edited by
Stefan Borgwardt, Universität Basel, Switzerland
Reviewed by
Bertus F. Jeronimus, University of Groningen, Netherlands; Stefan Weinmann, Vivantes Klinikum, Germany; Glen Spielmans, Metropolitan State University, United States
Updates

Check for updates
Copyright
© 2018 Hengartner and Plöderl.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.
*Correspondence: Michael P. Hengartner michaelpascal.hengartner@zhaw.ch
This article was submitted to Public Mental Health, a section of the journal Frontiers in Psychiatry
Disclaimer
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.