Skip to main content
Since 2004, revealing what drives you!

Proof that burnout science has been, at minimum, incompetent for 50 years.

They cannot decently plead ignorance that they are aggregating apples and oranges, then presenting the result as though every study had measured only apples. Here is the argued demonstration, using this 2019 meta-analysis by Koutsimani et al., which concentrates the majority of the field's problems.

The dossier's introduction states the central problem: "The Faking of Burnout was built precisely on a culture of the summary and the skim: the conclusion is cited without the concept, the logic, or the calculation ever being examined."

This meta-analysis is the archetype of it, because it checks every box of what I would call here scientific laundering: an operation by which the prestige of the method covers over the absence of proof, then hands back as certainty what the data never established.

The prestige of the form.

This is a systematic review with meta-analysis, registered on PROSPERO and conducted according to PRISMA — in other words, everything that sits at the top of the contemporary evidence hierarchy. For a hurried reader, the form already stands as guarantee. The methodological armor looks impeccable; the conclusion can then circulate without anyone going back to check whether the calculations actually contain it.

The illusion of an open question.

The authors ask whether burnout is conflated with depression or anxiety. But the separation is already written into the design: the studies selected enter the analysis with scores for "burnout," "depression," and "anxiety" already constituted as distinct variables. The corpus is dominated by the MBI, and when the results fail to settle the question, the discussion falls back on Maslach and Leiter to support the construct's autonomy — and therefore its validity and its distinctness.

The conclusive leap.

The calculations establish extreme heterogeneity, coefficients that shift depending on the instrument, and results that vary even more between certain populations. No discriminant-validity test demonstrates that the constructs are distinct. The abstract nonetheless pronounces "different and robust" objects.

This crossing of the evidentiary line then presents itself as an authoritative interpretation, one supposedly made possible by deep knowledge of the field and a fine command of statistics enabling the interpretation of results. The authors' presumed competence thus becomes the license to conclude beyond what their own analyses establish. The calculations produce instability; scientific authority reinterprets it as stability.

The transformation of a verdict into settled proof.

Once pronounced by a registered meta-analysis, PRISMA-compliant and published in a recognized journal, this conclusion changes status. It no longer appears as an interpretation the results cannot support, but as the output of the highest form of scientific synthesis. It can now be cited without carrying along the heterogeneity, the instrument-dependence, the incompatible populations, and the absence of any decision rule.

Adding this exhibit to the dossier alongside the other central meta-analyses is useful, because it exposes, in almost caricatural form, the mechanism by which the concept survives.

After forty-five years of research, the distinction between burnout, depression, and anxiety is still not demonstrated. It is presupposed by the instruments, weakened by the results, then restored in the conclusion. If this boundary cannot be scientifically established, the entire apparatus built on it — questionnaires, categories, prevention programs, expertise, and specialized interventions — loses the proof of its own autonomy.

This is what makes it archetypal: burnout does not hold up through proof, but through the repetition of its own claimed autonomy, in increasingly sophisticated methodological forms.



Here is the exhibit that has just been added to the FAKING BURNOUT V3 dossier:

Closing exhibit of Act III — Koutsimani et al. (2019): Instability Laundered into Proof of Distinction

Source: Koutsimani, Panagiota, Montgomery, Anthony, & Georganta, Katerina. "The Relationship Between Burnout, Depression, and Anxiety: A Systematic Review and Meta-Analysis." Frontiers in Psychology, Vol. 10, Article 284, March 13, 2019.

Status: Systematic review with meta-analysis, registered protocol (PROSPERO CRD42018090505), conducted according to the PRISMA standard. This format of quantitative synthesis is generally placed at the top of evidence hierarchies, on condition that the question asked, the studies selected, and the aggregation operation actually allow it to be answered. It is precisely this condition that the present exhibit examines. Published open access under a CC BY license. Corpus: 67 articles yielding 69 analysis entries for the burnout–depression relationship; 34 articles yielding 36 entries for the burnout–anxiety relationship. Since the same samples can produce several coefficients or appear under several instruments, no total of unique participants is advanced in this exhibit (the reason is set out in Section 7). The document explicitly takes a position in the debate over overlap between burnout and depression — a debate driven in particular by Bianchi, Schonfeld, and Laurent — and concludes in favor of the distinctness of the constructs, without setting out any decision rule capable of producing that verdict.

Notes on the instruments cited: Three instruments recur throughout this exhibit as the dominant measures of their respective corpora. They are introduced here once, to avoid weighing down each section with a repeated gloss.

MBI — Maslach Burnout Inventory. The most-used instrument in the study's corpus for measuring burnout: 55.1% of the studies on the burnout–depression relationship, 63.9% of those on burnout–anxiety.

PHQ — Patient Health Questionnaire. A self-report questionnaire for depression designed by Spitzer and colleagues, widely used in general medicine. The most-used instrument in the corpus for measuring depression: 36.1% of the relevant studies.

HADS — Hospital Anxiety and Depression Scale. A self-report scale developed by Zigmond and Snaith, designed to measure anxiety and depression jointly in non-psychiatric clinical settings. The most-used instrument in the corpus for measuring anxiety: 30.6% of the relevant studies.

In both relationships studied by the meta-analysis, the MBI occupies, on the burnout side, the same dominant position that the PHQ occupies on the depression side and that the HADS occupies on the anxiety side. This symmetry of status is what makes comparable the gaps produced by replacing each of these instruments, developed in Section 6.

Preamble — A Separation Already Made

The article situates itself explicitly within a continuity of prior work. Its introduction reviews studies concluding in favor of overlap and those concluding in favor of distinction, then notes that earlier results remain "inconclusive." In the discussion, it then invokes Maslach and Leiter and the "accumulated evidence" to support the claim that burnout and depression do not constitute a single illness.

Prior work thus provides the authors with an argumentative orientation. It does not provide their calculation with a decision rule.

More fundamentally, the meta-analysis operationally presupposes the very separation it claims to examine.

It selects studies in which:

  • one instrument measures "burnout";
  • another measures depression or anxiety;
  • the relationship between the two scores is then calculated.

In other words, each study enters the meta-analysis with two columns already named:

  • burnout score;
  • depression or anxiety score.

The aggregation never goes back upstream to ask whether the first column actually corresponds to an autonomous construct. It inherits that autonomy from the instruments and from the prior literature.

The meta-analysis therefore does not test the initial separation. It aggregates studies that had already put it into practice.

It then uses the degree of correlation between the two columns as though this could retroactively validate their separation.

This continuity of the field operates on three levels:

Conceptual continuity: burnout is introduced as an already-constituted syndrome, drawn principally from Maslach's dominant definition.

Instrumental continuity: the selected studies have already separated burnout, depression, and anxiety within their own questionnaires.

Continuity of authority: when the result is not enough to establish the distinction, the discussion cites Maslach and Leiter and their "accumulated evidence."

The analysis thus begins only after the separation has already been made.

1. Six Statements from a Single Document

The six passages below come from the same text. What their coexistence entails is established in the points that follow.

"Our findings revealed no conclusive overlap between burnout and depression and burnout and anxiety, indicating that they are different and robust constructs." (p. 1)
"Cohen provided rules of thumb for interpreting these effect sizes, suggesting that an r of 0.10, represents a 'small' effect size, 0.30 represents a 'medium' effect size and 0.50 represents a 'large' effect size (Cohen, 1992). However, researchers have suggested that the indiscriminate use of Cohen's generic small, medium, and large effect size values to characterize effect sizes in domains in which normative values do not apply is inappropriate and misleading (Lipsey et al., 2012)." (p. 5)
"Overall results indicated a significant effect (r = 0.520, SE = 0.012, 95% CI = 0.492, 0.547)." (p. 8)
"Although burnout and depression are associated with each other, the effect size is not so strong that it would suggest they are the same construct." (p. 14)
"Overall, our results suggest that the studies that used measures other than the MBI burnout tool could potentially be artificially inflating the association between burnout and depression and burnout and anxiety as well." (p. 14)
"According to our results, it is possible that individuals who are more prone to experiencing higher levels of anxiety (trait anxiety) are also more likely to develop burnout as well." (p. 14)

2. The Founding Inversion: Indeterminacy Turned into a Verdict

The first slippage precedes the interpretation of a single coefficient. The abstract states the absence of a "conclusive" overlap ("no conclusive overlap"), then infers from it that the constructs are "different and robust." But the absence of conclusive proof of an overlap is not conclusive proof of its absence. It leaves the question open — nothing more. It would take a further study to conclude that no link exists. Here, indeterminacy is turned into a verdict of distinction.

The present examination does not dispute that a hypothesis of distinction may be provisionally retained while the question stays open. It disputes that this holding pattern is presented as a result produced by the analysis. A failure to demonstrate overlap may justify withholding a conclusion; it does not license writing that the results indicate different and robust constructs.

And the verdict is doubled. "Different" and "robust" are two separate claims, neither of which follows from the analyses or the data. Even a correct demonstration of distinction — which, as the following sections show, never took place — would still leave the question of robustness entirely open: a construct can be distinct from another while remaining unstable, poorly defined, and dependent on its instrument. The abstract thus asserts twice as much as the calculations permit.

"Overall, according to our results burnout and depression and burnout and anxiety appear to be different constructs that share some common characteristics and they probably develop in tandem, rather they fall into the same category with different names being used to describe them." (p. 14)
"Nevertheless, according to our meta-analysis results, burnout and depression and burnout and anxiety appear to be different rather the same constructs." (p. 14)

In the discussion (p. 14), the authors also call for future studies to examine the psychosocial and neurobiological bases of these constructs. In their conclusion (p. 15), they acknowledge that further work remains necessary to clarify their relationships and reach safer conclusions.

These precautions exist. They do not, however, repair the design. The word "appear" signals a reservation, but no analysis in the document establishes why distinction would be better supported than overlap. Likewise, the claim that the phenomena "probably" develop in tandem corresponds to no probability computed by the model.

The document thus operates under two regimes. In the body of the text, distinction is presented with hedges — "appear," "probably" — and accompanied by a call for further research. In the abstract, these reservations disappear: the results are said to indicate "different and robust" constructs. The caution in the body of the text does not erase the categorical verdict published in the summary; it shows, on the contrary, that the text itself recognizes an uncertainty its abstract does not retain.

3. A Measured Association, a Distinction Without a Decision Rule

To follow the operation, one reminder suffices. A correlation coefficient theoretically ranges from −1 to +1. A value near 0 indicates that the two series of scores do not vary together linearly; a positive value indicates that they tend to increase together; a value near +1 indicates a near-perfect linear relationship. Even a coefficient of +1 would not, by itself, demonstrate that two questionnaires measure conceptually the same object: it would only establish that their scores move together perfectly. The document's method section installs the conventional reading grid — 0.50 signals a "large" effect — then hedges it with a warning: these generic thresholds do not apply mechanically, interpretation requires contextual grounding (both passages are quoted in Section 1). This double requirement calls for a criterion: if Cohen's thresholds are not enough, the document must supply the contextual criterion that replaces them — and state it before the result.

The result comes out at 0.520. It crosses the conventional threshold for a large effect stated in that same method section, without any criterion. To be precise about the charge: Cohen's threshold is not a criterion of identity between constructs, and a "large" effect does not demonstrate that two questionnaires measure the same thing — this exhibit claims no such thing. The real problem lies elsewhere, and it is more serious. The authors substitute no device for the generic thresholds that would allow a decision on identity or distinction. Having obtained an association their own grid calls large, the discussion simply declares it "not so strong." Not strong enough compared to what? No numerical threshold, no model, no test set in advance determines the value at which distinction could be retained. The text never says.

As a result, the rule leading to the conclusion is not reproducible. Another team, faced with exactly the same numbers, could just as legitimately call the association "strong enough," since no criterion allows the matter to be settled. The coefficient is reproducible; the verdict attached to it is not. A conclusion that depends on the judgment of whoever states it no longer follows from the calculation — it belongs to opinion. The sequence exposes the real order of operations: the result first, the bar afterward, adjusted to the conclusion ultimately published. Distinction is added after the calculation.

4. The Missing Apparatus: A Question of Overlap, an Answer of Association

The stated question concerns the overlap of constructs: are burnout and depression the same thing? The analysis actually carried out measures something else: the degree to which two sets of questionnaire scores vary together. The article thus changes questions mid-operation — it poses a question of kind, it computes a degree of association.

A correlation, even correctly calculated, tells us directly nothing about any of the elements the answer actually depends on: not diagnostic overlap (would the same individuals be identified on both sides?), not factorial identity (do the items load on the same underlying dimension?), not the interchangeability of instruments, not the existence of a shared latent variable, not discriminant validity. Establishing that one construct is distinct from another would have required a device specifically built for that question — latent measurement models, for instance, comparison of competing models (a single common factor against two correlated factors), correction for attenuation due to instrument unreliability, or distinction criteria set before the calculation. The document performs none of these operations. It aggregates correlations between questionnaires, then pronounces a verdict on the nature of the objects.

A discriminant-validity threshold could not, in any case, be applied after the fact to the aggregated coefficient: such criteria concern specific instruments and specific measurement models, often built from the relationships between their items or their latent variables. The mean r produced here from a dozen heterogeneous instruments is not that statistic.

One point should be stated plainly for every reader, including non-specialists: two questionnaires measuring largely the same phenomenon have no particular reason to correlate at 1, and in practice, psychological instruments very rarely reach a perfect correlation. The reasons are ordinary — measurement error, differently worded items, partially distinct content, heterogeneous populations and contexts, restricted distributions, different testing occasions, imperfect instrument reliability. The gap between 0.52 and 1.00 contains all of this before it contains the slightest proof of distinction between the two constructs. The document's reasoning nonetheless turns the imperfect, expected character of the relationship between two measures into a property of the object: because the correlation is not perfect, the phenomena are said to be different. No test specifically devoted to that distinction ever intervenes between the two.

5. An Aggregate with No Homogeneous Object

The corpus aggregated under the word "burnout" gathers the MBI and its variants (MBI-HSS, MBI-GS), the CBI, the SMBM, the OLBI, Pines' Burnout Measure, the BSI, the ProQOL, the PFI, the HBI, the OBI, the SMBQ — along with total scores and isolated dimensions (emotional exhaustion, depersonalization, professional accomplishment), crossed against fourteen depression instruments and seventeen anxiety instruments (Table 3, p. 10). The r = 0.520 therefore does not measure the relationship between "burnout" and "depression": it averages relationships across a multitude of measurement operations that do not define burnout in the same way.

Heterogeneity quantifies this problem: I² = 98.432% for depression, I² = 95.367% for anxiety (pp. 8–9). This index estimates that more than 95% of the observed variability among results is not explained by sampling error alone, but by heterogeneity in the material — instruments, populations, dimensions used, and measurement operations. The document itself sets 75% as a "high" level of heterogeneity (pp. 5–6); both of its corpora exceed 95%. An I² of this magnitude does not by itself prove the construct does not exist; it invalidates the use of the mean as a simple representation of a common, stable relationship. The r = 0.520 does not, then, summarize a relationship consistently found from one study to the next. It is the mathematical average of results that differ almost entirely depending on the instruments, populations, dimensions used, and measurement operations. Presented alone, this average gives the appearance of a common, stable result where the corpus mainly documents the instability of the relationship being measured. It gathers under a single figure studies that do not give the same answer to the question being asked.

The moderator analyses explore this partially and discover precisely that it depends, among other things, on the instrument used. But this finding is never treated for what it is: a central result that directly undermines the general verdict. If the relationship between burnout and depression changes sharply depending on the tool chosen to measure burnout, the meta-analysis can no longer present its mean coefficient as a stable property of the constructs themselves. What varies is not a peripheral detail around a robust result. It is the result. The measure chosen participates in producing the degree of overlap observed.

The abstract nonetheless pronounces "robust" constructs, neutralizing exactly what the internal analyses have just established: that their degree of distinction depends on the operations that measure them. The conditions under which the result was produced are demoted to the rank of secondary variation, and the mean these conditions destabilize is then elevated to the rank of a property of the objects. Heterogeneity is thus not simply noted and then insufficiently discussed: it is explored, partly explained, and then set aside at precisely the moment it should block the verdict.

6. The Subgroups: Distinction Depends on the Instrument

The moderator analyses break down the overall result by instrument and by dimension used. They produce two distinct series: one for the burnout–depression relationship, one for the burnout–anxiety relationship.

Examining them reveals the same movement on both sides, at different magnitudes. For depression, instrument-dependence is massive and bilateral: it affects the measurement of burnout as much as the measurement of depression itself. For anxiety, the dependence is also bilateral, but narrower.

6.1. Burnout and Depression: Distinction Varies from Simple to Almost Double

The document's moderator analyses (pp. 9–10) produce, for the burnout–depression relationship, the following series:

Aggregated Measurement Operation Correlation with Depression
Whole corpus r = 0.520
Studies using the MBI r = 0.472
Studies using another instrument r = 0.622
Emotional exhaustion (MBI) r = 0.508
Other MBI dimensions r = 0.409
Non-MBI instruments, total score r = 0.749 (SE = 0.136)
Non-MBI instruments, subscales r = 0.608

These figures should be read very simply. The researchers claim to be determining whether burnout is distinct from depression. Yet the answer changes sharply depending on which questionnaire is used to represent burnout. Under some measurement operations, the scores appear relatively distant from each other; under others, they overlap massively. The choice of instrument therefore itself produces a substantial part of the answer, alongside any stable property of the phenomenon.

Two facts stand out. First fact: the level of overlap with depression varies from simple to almost double — from r = 0.409 to r = 0.749 — depending on the instrument, the dimension, and the scoring method used. Second fact: when burnout is treated as a total score by instruments competing with the MBI, overlap reaches r = 0.749 — roughly 56% of shared variance between the scores.

A reservation must accompany this last value. Its standard error — 0.136 — is by far the largest of all the subgroups, more than ten times that of the overall result, and the published text does not specify how many studies make up this subgroup. It is therefore probably small, and its estimate remains imprecise. The figure of 0.749 cannot be put forward as a stable value.

But this imprecision does not make the result disappear. The lower bound of its confidence interval — r = 0.643 — still sits above the level the document's own method calls "large." Even under the most unfavorable assumption compatible with the published data, non-MBI instruments using a total score document a substantial overlap with depression scores. Imprecision forbids pinning down the exact value; it does not undo the finding.

More importantly, the charge does not rest on this subgroup alone. Even setting r = 0.749 aside, the main contrast remains: studies using the MBI produce a mean correlation of r = 0.472, versus r = 0.622 for those using other instruments. The result's dependence on the choice of measure stands intact.

This instrument-dependence does not concern only how burnout is measured. The authors also tested the effect of the instrument used to measure depression:

  • studies using the PHQ: r = 0.628;
  • studies using other depression instruments: r = 0.481.

The gap is 0.147 — almost exactly the same magnitude as the gap produced by the choice of burnout instrument: r = 0.472 with the MBI versus r = 0.622 with other instruments, a gap of 0.150.

The burnout–depression relationship is thus bilaterally dependent on instruments:

  • when the burnout measure changes, the correlation changes;
  • when the depression measure changes, it changes almost as much.

This elementary demonstration is nonetheless never carried through to the discussion. When the depression questionnaire produces the highest correlation, the document does not interpret it. When burnout questionnaires competing with the MBI produce the highest correlation, the document instead suspects them of "artificially inflating" the association between burnout and depression (p. 14, quoted in Section 1).

Faced with two nearly identical variations, the discussion examines only one: the one that preserves the dominant instrument. The symmetrical reading is left out of the text.

The composition of the MBI subcorpus adds, alongside emotional exhaustion — r = 0.508 — coefficients drawn from depersonalization and professional accomplishment, grouped among the "other dimensions" and less correlated with depression — r = 0.409. Their inclusion can mechanically lower the mean association attributed to MBI studies, without that lowering demonstrating the existence of an autonomous syndrome. Diversifying the content of a measure can move its scores away from an external criterion; this distancing does not prove that a distinct phenomenon has been isolated.

The results thus open two competing interpretations: instruments other than the MBI may artificially inflate overlap, or the composition of the MBI subcorpus may instead dilute it. Since the meta-analysis contains no analysis capable of adjudicating between them, both readings must remain equally open.

The discussion retains only one: inflation produced by competing instruments. It thereby dismisses, without examining it, the symmetrical hypothesis of dilution produced by the MBI subcorpus — precisely the interpretation that would weaken the dominant instrument.

The choice made in the discussion thus results from no empirical arbitration contained in the meta-analysis. And a distinction that varies with the instrument cannot be attributed to the construct as a stable property: it also measures the effect of the tool that produced it. The discussion thus presents as a difference between burnout and depression what its own analyses show to depend on how burnout is measured. Its own results directly call into question the dominant instrument's capacity to support a stable distinction, without the document ever spelling out or analyzing that consequence.

In this first movement, instrumental variation constitutes the central result: the measure chosen does not reveal a pre-existing level of overlap — it constructs the very form in which that overlap becomes measurable, and can then be presented as a property of the constructs.

6.2. Burnout and Anxiety: The Same Dependence, on Both Sides of the Measure

The moderator analyses produce, for the burnout–anxiety relationship, the following series:

Aggregated Measurement Operation Correlation with Anxiety
Whole corpus r = 0.460
Studies using the MBI r = 0.451
Studies using another instrument r = 0.482
Emotional exhaustion (MBI) r = 0.472
Other MBI dimensions r = 0.426
Non-MBI instruments, total score r = 0.494 (SE = 0.060)
Non-MBI instruments, subscales r = 0.499 (SE = 0.052)

The spread is less dramatic than for depression. It does not run from r = 0.409 to r = 0.749, but from r = 0.426 to r = 0.499, depending on the dimensions and scoring methods used.

This difference in magnitude should not lead one to dismiss the result. The same regularities appear. Studies using the MBI produce a slightly lower mean association — r = 0.451 — than studies using other instruments — r = 0.482. Within the MBI, emotional exhaustion is more strongly associated with anxiety — r = 0.472 — than the other dimensions — r = 0.426. Finally, non-MBI instruments producing a total score or subscales reach the highest values — r = 0.494 and r = 0.499 respectively.

The two highest non-MBI estimates should be read with caution. Their standard errors — 0.060 for total scores and 0.052 for subscales — are noticeably larger than that of the overall result. They indicate a direction, but their exact value cannot be treated as stable.

Instrument-dependence does not, however, rest on these two subgroups alone. Even setting them aside, studies using the MBI produce a weaker mean association with anxiety — r = 0.451 — than those using other instruments — r = 0.482. Within the MBI, emotional exhaustion reaches r = 0.472, against r = 0.426 for the other dimensions.

The same shift thus appears at several levels: the association increases when exhaustion is isolated or when burnout is measured with another instrument; it decreases when the other MBI dimensions are folded in. The questionnaire chosen to construct the burnout score directly shapes the level of overlap produced with anxiety.

The coefficient thus already varies according to the instrument used to measure burnout. It also varies when the instrument used to measure anxiety changes:

Instrument Used to Measure Anxiety Correlation with Burnout
HADS r = 0.507
Other instruments r = 0.437

The correlation thus moves from r = 0.437 to r = 0.507 depending on the instrument chosen to represent anxiety. For a supposedly identical phenomenon, the level of overlap also changes when the researcher swaps out the tool on the other side of the equation.

The authors themselves acknowledge part of this problem:

"However, it should be noted that there was variability in the inventories that were used in the research studies for assessing anxiety. It is possible that the effect sizes between burnout and anxiety would differ if there was a common widely used tool for assessing anxiety." (p. 14)

This admission is both partial and, at the same time, more decisive than a mere acknowledgment of instrumental diversity. The authors concede that effect sizes might be different if all the studies used the same tool to measure anxiety. The aggregated coefficient thus gathers results produced with different questionnaires, which do not necessarily define anxiety the same way, drawn from different populations and under different conditions. The meta-analysis nonetheless treats these coefficients as comparable estimates of a single relationship between burnout and anxiety. It pools them into a general mean even as the authors themselves acknowledge that this mean could change if one of the measurement operations were standardized. The r = 0.460 therefore does not summarize an identical relationship recovered across the studies: it aggregates relationships built under different instrumental configurations, then interprets their mean as though they all came from a single measurement operation.

They know they are aggregating apples and oranges, then present the result as though every study had measured only apples.

This acknowledgment, however, concerns only the instrument used to measure anxiety. The moderator analyses establish that the choice of burnout instrument also shifts the correlation: it reaches r = 0.451 in studies using the MBI, against r = 0.482 in those using other instruments; within the MBI, it moves from r = 0.426 for the other dimensions to r = 0.472 when emotional exhaustion is isolated.

Both sides of the relationship thus vary with the tools that construct them. The correlation changes when the anxiety measure is swapped, then changes again when the burnout measure is swapped. It cannot, then, be attributed to the phenomena alone as a stable property of their relationship: it also expresses the particular combination of instruments, dimensions, populations, and conditions gathered in each study.

The relationship with depression shows a massive bilateral dependence: the coefficient shifts almost as much when the depression measure changes as when the burnout measure changes. In the case of anxiety, the gaps are narrower, but the same dependence appears on both sides of the relationship — and the authors themselves acknowledge that using a common tool could change the effect sizes.

This double dependence makes the verdict of robustness contradict the study's own findings. The authors call "different and robust" constructs whose own analyses show that the degree of distinction shifts with the burnout instrument, and could shift further still if the anxiety measure were standardized. They conclude in favor of stability from results that establish instability.

6.3. Overlap Also Varies with the Population Studied

The first two movements show that the result changes depending on the instrument chosen — on one side for measuring burnout, on the other for measuring depression or anxiety. A third factor produces an even wider gap: the occupational group studied.

The document's moderator analyses (p. 9) produce, for the burnout–depression relationship, the following series by occupational status:

Population Studied Correlation with Depression
Educational staff r = 0.679 (SE = 0.049)
Healthcare staff r = 0.495
General employed population r = 0.449

The gap between the top and bottom of this series reaches 0.230 — more than the gap produced by the choice of burnout instrument (0.150) and by the choice of depression instrument (0.147). Among all the published moderator analyses, the gap between occupational groups is the largest of all. This does not prove that occupation alone produces this variation, but it shows that the overall r varies more between the populations studied than between the instruments compared.

The lower bound of the confidence interval for educational staff — r = 0.609 — still sits above the level the document's method conventionally calls "large." Even accounting for the uncertainty around the estimate, the value therefore remains large by the authors' own grid. This does not rule out other possible differences between the studies, but it forbids treating r = 0.679 as the mere product of an overly imprecise estimate.

A caveat is needed before going further. The article does not test a model that would isolate the effect of occupation while holding instrument, quality, and country constant. Studies conducted among teachers may differ from the others on many grounds beyond occupation alone — the questionnaire used, the national context, the year of publication. One cannot therefore write that occupational status explains or causes this gap. Among teachers, burnout and depression scores are far more tightly linked than in the general population. This does not prove they measure exactly the same thing, but it makes it much harder to present them as two clearly separate realities.

By folding these two results into a single mean, the r = 0.520 erases this difference. It gives the appearance of a general relationship while actually describing neither group.

On the anxiety side, the comparison between groups shows a narrower gap: healthcare staff obtain r = 0.436, against r = 0.492 for the general employed population. The educational subgroup could not be tested — only two studies involved teachers, a number the authors themselves judged insufficient to produce an estimate.

The three movements now converge on the same finding. The coefficient obtained varies when the burnout measure is swapped. It varies almost as much when the depression measure is swapped. It varies even more across occupational subcorpora.

The article does not allow us to determine how much of these gaps is separately attributable to instruments, populations, or other study characteristics. But it establishes one thing: no stable value holds across these different configurations.

The r = 0.520 cannot, therefore, be presented as the general measure of the relationship between burnout and depression. It constitutes the mean of relationships produced with different instruments, among different populations, and under different conditions — and it is precisely this heterogeneous mean that the abstract declares "robust." Robustness is now undermined not only by the diversity of instruments, but by the instability of the result across every configuration examined.

7. The Non-Independence of the Coefficients

The forest plots (Figures 3 and 4) line up several coefficients drawn from the same studies: Favrod et al. (2018) appears in entries (a) through (i), Cardozo et al. (2012) in several entries, and many studies contribute separate lines by dimension (DEP, EX, PA) calculated on the same participants. The corpus tables show the same phenomenon: Bianchi & Laurent (2015) appears twice with the same 54 participants, one entry per burnout instrument; Choi et al. (2018) appears twice with the same participants. Adding up the rows would therefore count the same people more than once — which is why this exhibit puts forward no total participant count.

The analysis section describes a standard random-effects model (pp. 4–5). The text documents no procedure for correcting the dependence between coefficients drawn from the same samples: no multilevel model, no robust variance estimation, no prior aggregation by study. What the article presents as 69 studies therefore does not necessarily correspond to 69 independent pieces of statistical information.

The effect of such treatment, when several correlated coefficients are treated as independent observations, is well known: studies contributing more dimensions or measures are weighted more heavily than they should be, and confidence intervals narrow beyond what the material can support. The standard error of 0.012 displayed alongside the main result then gives the appearance of considerable precision for a corpus whose heterogeneity exceeds 98%. Without access to the analysis data, recalculation lies beyond the scope of this exhibit; the absence of any documented procedure, however, is a fact of the text.

8. Quality, and the Drift of the Question

The quality assessment relies on the Quality Assessment Tool for Observational Cohort and Cross-Sectional Studies — fourteen criteria covering, among other things, recruitment, observation methods, exposure, and follow-up (p. 4). Studies classified as "good quality" produce lower correlations — r = 0.488, against r = 0.565 for "fair quality" studies (p. 10).

This quality ranking, however, does not answer the question the article poses. The grid evaluates how the studies were conducted; it does not evaluate what their instruments measure. None of its criteria test the content validity of burnout, the existence of a diagnostic standard, or its discriminant validity against depression.

The lower coefficient obtained in better-rated studies establishes only one thing: those studies produce a lower association. To infer two distinct constructs from this, one would still need to demonstrate that the instruments actually measure two distinct objects. The abstract skips this step and turns the quality of observation into proof about the nature of the object observed.

9. Causality Ruled Out, Then Reintroduced

The depression corpus is 87% cross-sectional; the anxiety corpus, 97.2%, with a single longitudinal study (p. 8). A cross-sectional study photographs a single moment: the same people complete both questionnaires at the same time. The text contains the corresponding admission:

"Therefore, although the burnout—depression and the burnout—anxiety relationships are found to be related, we are still not able to know whether these relationships are casual [sic]." (p. 14) — [The typo — "casual" for "causal" — appears in the published text.]

The same page nonetheless advances the proposition quoted in Section 1: individuals with a durable anxious disposition — "trait anxiety" — would be more likely to develop burnout. (State anxiety refers to a momentary, situation-linked anxiety; trait anxiety refers to a relatively durable disposition toward anxiety.) This proposition does not follow from the data. A cross-sectional correlation cannot determine: whether anxiety precedes the burnout score; whether exhaustion increases anxiety; whether the two scores measure a common distress; or whether a third variable produces both simultaneously.

Four readings remain compatible with these same numbers. Only one appears in the discussion: the one placing a durable anxious disposition upstream of burnout, drawn from data incapable of establishing that causal order.

The document gives no way to know whether the other possibilities went unseen, were considered and set aside, or were deliberately dropped. Their disappearance nonetheless produces an effect this dossier must examine. By making anxiety a predisposition to burnout, the discussion presupposes that the two are already distinct phenomena: one becomes the cause, the other its result.

This reading thus turns an indeterminate relationship — or non-relationship — between scores into a causal sequence between two autonomous objects. It simultaneously rules out the hypothesis of a shared or identical distress, that of an overlap produced by the instruments, and that of a burnout not sufficiently distinct from anxiety or depression to constitute its own object. And this is precisely what has occupied the field since its beginnings: attempting to establish that burnout constitutes its own object, without ever managing to demonstrate it.

The reasoning is circular: the discussion presupposes burnout's autonomy in order to construct a causal sequence, and that sequence then helps confirm the autonomy it presupposed. Yet cross-sectional data can establish neither the causal order nor the existence of two distinct constructs. The discussion thus turns two absences of proof into confirmation of the concept.

10. The Verdict Imported from the Holders of the Dominant Definition

The discussion's final ruling rests on two supports:

"Burnout is an occupationally-specific dysphoria that is distinct from depression as a broadly based mental illness (Maslach et al., 2001)." (p. 14)
"Maslach and Leiter (2016) have argued that while studies confirm that burnout and depression are not independent, claiming that they are simply the same mental illness is not supported by the accumulated evidence." (p. 14)

Yet the document's analyses do not demonstrate the distinction (Sections 2 through 8); the discussion ultimately grounds it in Maslach et al. (2001) and Maslach & Leiter (2016); the MBI dominates the corpus — 55.1% of depression studies, 63.9% of anxiety studies; Maslach holds the copyright on the MBI-HSS, and Maslach and Leiter are co-authors of the MBI-GS.

Structural inference: by handing the final ruling to these sources, the document hands the arbitration of the concept over to actors with a direct interest — scientific and instrumental — in its continuity. The distinction of burnout from depression supports the dominant instrument's own value: a measure specifically devoted to burnout presupposes that it captures something a depression measure does not already capture. In other words, at precisely the point where the data do not settle the matter, the discussion stops demonstrating anything and instead cites those who fixed the dominant definition.

Where the calculations fail to produce the distinction, it is the discussion that imports it from the authors of the dominant tool. This reimportation of authority should be traced back to the field's editorial and commercial infrastructure.

One further element deserves attention, since it both stages and completes this meta-analysis's configuration. Its introduction reproduces the canonical origin story of burnout that Maslach and Schaufeli anchored:

"In particular, two independent researchers, Herbert Freudenberger, a psychiatrist, and Christina Maslach, a social psychologist, were the first researchers who began examining burnout." (p. 2)

This dossier establishes, from dated primary texts, that this claim to priority does not describe the concept's actual history: it erases prior work and retrospectively reconstructs Freudenberger and Maslach as its founders.

The meta-analysis thus does not simply study an already-available construct. Before even examining the data, it adopts the narrative that founds its autonomy, installs Maslach as the originating authority, and anchors the analysis in a continuity. It then aggregates a corpus dominated by the MBI, and invokes Maslach and Leiter in the discussion when its own calculations fail to establish the distinction.

The same authority thus intervenes at every stage: it historically founds the object, supplies its dominant instrument, and finally underwrites its theoretical autonomy. Yet the article presents itself as an independent synthesis tasked with arbitrating the overlap question. It nonetheless enters the question through Maslach's narrative, deals mainly with results produced by her instrument, and exits through Maslach's verdict.

Conclusion

This meta-analysis does not demonstrate that burnout constitutes an object distinct from depression or anxiety. It presupposes that distinction at the outset, aggregates studies that had already put it into practice in their questionnaires, then presents their mean as confirmation of what it had already assumed before any calculation.

The examination establishes the following facts:

  • The stated question is never tested. No discriminant-validity test, no decision threshold, and no model allow one to determine at what result two constructs could be declared distinct. The authors calculate correlations between scores; they do not establish the nature of the objects those scores are supposed to represent.
  • The overall mean gathers studies that do not measure the same thing in the same way. Instruments, dimensions, populations, and study conditions differ sharply. Heterogeneity is nearly total. The r = 0.520 does not summarize a stable relationship found from study to study: it melts, into a single figure, results that do not tell the same story.
  • The result changes when the burnout measure is replaced. The correlation with depression reaches r = 0.472 with the MBI, versus r = 0.622 with other instruments. Depending on the dimensions and calculation methods used, it ranges from r = 0.409 to r = 0.749. The measure does not, then, reveal an already-formed level of distinction: it helps produce it.
  • The result changes almost as much when the depression measure is replaced. The correlation reaches r = 0.628 with the PHQ, versus r = 0.481 with other questionnaires. The gap produced by the depression measure — 0.147 — is nearly identical to that produced by the burnout measure — 0.150. The coefficient is stable on neither side of the relationship.
  • The same dependence appears for anxiety. The correlation varies with the burnout measure — r = 0.451 with the MBI versus r = 0.482 with other instruments — and varies again with the anxiety measure — r = 0.507 with the HADS versus r = 0.437 with other questionnaires. The authors themselves acknowledge that effect sizes could be different if all studies used the same tool. They therefore know the instruments do not necessarily produce the same relationship, yet they aggregate their results and reason as though every study had measured the same thing.
  • The mean shifts further still between the populations studied. Overlap with depression reaches r = 0.679 in the subcorpus devoted to educational staff, against r = 0.449 in the subcorpus of the general employed population. The gap reaches 0.230, larger than the gaps tied to instruments. The article does not allow this difference to be attributed to occupation alone, but it establishes that the r = 0.520 truly describes neither group: it erases their difference within a single mean.
  • The precision displayed around the main result is not guaranteed by the actual structure of the data. Several studies contribute to the calculation more than once — with different instruments, different dimensions, or several coefficients drawn from the same participants. These results are therefore not fully independent of one another. Yet no documented correction accounts for this repetition. The r = 0.520 comes with a strikingly small standard error — SE = 0.012 — which gives the appearance of an extremely precise result, even though much of the aggregated data comes from the same studies, and sometimes the same people.
  • Methodological quality is diverted from what it actually assesses. The grid used judges recruitment, observation, exposure, and follow-up. It tests neither the validity of burnout nor its distinction from depression. The abstract nonetheless turns a weaker correlation in better-rated studies into an argument for two distinct constructs. The quality of observation becomes, with no intervening test, proof about the nature of the object observed.
  • Causality is declared impossible, then reintroduced in the discussion. The corpus is almost entirely cross-sectional. The authors acknowledge they cannot determine whether anxiety precedes burnout, whether exhaustion increases anxiety, whether the two scores measure a single distress, or whether a third variable produces both. Only one reading nonetheless appears in the discussion: that a durable anxious disposition favors the development of burnout. This reading turns a directionless relationship into a causal sequence between two objects already assumed to be autonomous. Nothing in the data establishes that causal order, the distinction of the constructs, or the autonomy of burnout.
  • When the calculations fail to produce the expected distinction, the discussion imports it from Maslach. The article enters the question through the genealogical narrative that installs Maslach as founder, deals mainly with results produced by her instrument, then invokes Maslach and Leiter to support the claim that burnout and depression do not constitute a single illness. The same authority founds the object historically, supplies its dominant instrument, and finally underwrites its theoretical autonomy.

This meta-analysis establishes that the result depends on the instruments used, the dimensions retained, the populations aggregated, and the conditions under which the scores were produced. It discovers no stable value that could be attributed to burnout as a general property. It demonstrates the instability of the measurement, then attributes stability to the construct.

It thereby transforms:

  • the absence of a distinction test into a verdict of distinction;
  • the heterogeneity of the results into a general mean;
  • dependence on instruments into a property of the phenomena;
  • the impossibility of establishing causality into a causal story;
  • the absence of proof of autonomy into confirmation of burnout's autonomy.

It is not, then, the phenomenon that is distinguished by the measurement: it is the operations of measurement and aggregation that produce the level of distinction subsequently attributed to the phenomenon.

The sequence is complete: the authors posit two already-separated objects, aggregate scores that vary with instruments and populations, run no test capable of establishing their distinction, note extreme heterogeneity, then turn this indeterminacy into a certain conclusion. When their calculations no longer contain the verdict they were looking for, the discussion retrieves it from the authority that historically defined the object and supplied its dominant instrument.

This exhibit's findings can be summarized in three points:

  • The result is not stable. It changes when the instruments used to measure burnout, depression, or anxiety change, and varies even more across certain populations studied. The overall coefficient describes no common relationship found across the corpus: it absorbs different results into a single mean.
  • The conclusion is not produced by the analyses. No test, no threshold set in advance, and no decision rule allow one to move from the calculated correlations to the claim of distinct, robust constructs. The results leave several interpretations open; the abstract closes them into a verdict.
  • Burnout's autonomy is produced by circular reasoning. The construct is separated from depression and anxiety in the instruments before ever being examined; this prior separation is then aggregated, and presented as its own confirmation. When the calculations are not enough to establish the distinction, the discussion retrieves it from the authority that defined the object and supplied its dominant instrument.

This exhibit does not merely show that Koutsimani et al. fail to demonstrate burnout's autonomy. It shows how a presupposed distinction — unstable in the results and absent from the calculations — can nonetheless be turned into scientific certainty, and then circulate through the field as settled proof.

A registered meta-analysis, conducted according to PRISMA and published open access, thus pronounces an ontological verdict — "different and robust constructs" — that none of its calculations contain.

This exhibit does not offer a fragile proof of burnout's autonomy. It documents an absence of proof transformed into scientific certainty.

The field cites this verdict.
The verdict cites Maslach.

"Excellence is the result of consistent improvement."

Philippe Vivier

©

Philippevivier.com. All rights reserved.

Article L122-4 of the Code of Intellectual Property: "Any representation or reproduction in whole or in part without the consent of the author [...] is illegal. The same applies to translation, adaptation or transformation, arrangement or reproduction by any art or process."

History & Infos


Practice founded in 2004.
Website and content redesigned in 2012.
SIRET NUMBER: 48990345000091

Legal information.


Addresses


  • 254 rue lecourbe
    75015 Paris
  • 23 avenue de coulaoun
    64200 Biarritz
  • 71 allée de terre vieille
    33160 St Médard en Jalles
  • 16 Pl. des Quinconces
    33000 Bordeaux

Contact