Subjects · Psychology & Behavioral Sciences
Psychology & Behavioral Sciences: Correlation-causation mapping
Catch the exact sentence where a correlational finding quietly becomes a causal claim, which is the field's single most common public-facing error, and it always arrives wearing a plausible story.
What you'll be able to do: Catch the exact sentence where a correlational finding quietly becomes a causal claim, which is the field's single most common public-facing error, and it always arrives wearing a plausible story.
Scope note. Cause-consequence mapping also anchors history's cause-consequence mapping (article 01 there) and the natural sciences' version (article 07 there), both about building a causal web from a set of events or a practical result. Psychology's version is narrower and more specific: it's almost always about one pair of variables and one unstated third factor, and the failure mode isn't building the wrong web, it's skipping the web-building step entirely and asserting an arrow that the design can't support.
The sentence where it happens
"Students who eat breakfast perform better in school" becomes, one sentence later, "so make sure children eat breakfast." "People who meditate report lower stress" becomes "so meditation reduces stress." "Teenagers who spend more time on social media report more anxiety" becomes "social media causes anxiety in teenagers."
Every one of those first halves can be true and well-supported. The join to the second half is where the argument actually happens, and in an enormous amount of psychology writing, popular and academic alike, the join isn't argued. It's just typed, because a correlation with a plausible causal story attached feels like a causal finding, and psychology is unusually rich in plausible stories, because the variables are all things people already have opinions about.
This is the field's signature error, more than any other single mistake, and it survives in students who can recite "correlation isn't causation" as a rule and then commit it two paragraphs later, because reciting the rule and catching your own next sentence doing exactly the thing it warns against are different skills.
The three culprits, by name
Third variable. Some unmeasured factor causes both. Family socioeconomic status plausibly drives both breakfast-eating and school performance, without breakfast doing anything to performance directly, this is a large part of why the marshmallow test's original predictive power shrank once studies controlled for family background (article 01).
Reverse causation. The arrow the story assumes might run backwards. People who are already less stressed may be more likely to take up and stick with meditation, rather than meditation making them less stressed, plausible in either direction, and cross-sectional data can't tell you which.
Selection effects. People aren't randomly assigned to the groups being compared, and whatever made them end up in one group rather than the other may be the real driver. Teenagers who are already more anxious might use social media differently (more, less, or in different ways) than teenagers who aren't, which would produce the correlation without social media use causing anything.
The drill
Here's a finding: "[claim]." Map the causal claim explicitly: what's the correlation actually reported, and what would have to be true for it to be causal rather than one of these three, a third variable causing both, reverse causation, or a selection effect? Which of the three is most plausible here, and does the study's actual design (correlational, longitudinal, or experimental with random assignment) let it rule any of them out?
The design question at the end is the one that resolves it. A true experiment with random assignment to conditions can rule out selection effects and reverse causation by construction, that's what random assignment buys you. A correlational or cross-sectional study can't, no matter how large the sample or how significant the result (article 04).
Applying it to yourself, where it's hardest
The method gets harder exactly where it matters most: claims about your own life.
I read that [habit/trait] is linked to [outcome], and I do [habit/trait]. Before I take that as a reason to change anything, map the causal claim: third variable, reverse causation, or selection effect, which is most plausible for this specific finding, and what would the study need to have done to rule it out?
This is where motivated reasoning is strongest, because a correlational finding that flatters a habit you already have is much easier to accept uncritically than one that doesn't, the subject's reader-directed problem (see the primary article) showing up in miniature.
Pitfalls
- Stopping at "correlation isn't causation" as a slogan. The skill is naming which of the three specific mechanisms applies, not reciting the rule.
- Assuming a plausible story rules out the alternatives. Plausibility is not evidence; it's the reason the error is hard to catch.
- Missing that longitudinal data narrows but doesn't close the gap. Knowing X came before Y helps rule out simple reverse causation but not a third variable that produces both on a delay.
- Applying the check to other people's beliefs and not your own. The habit is worth the most exactly where you have a stake in the answer.
- The tell: you can state the rule correctly on a methods question and still write "so you should" two sentences after citing a correlational study.
Try this today
Find a psychology claim you've recently accepted without examining it, ideally one that's about a habit or trait of your own. Map it: what's the correlation, which of the three mechanisms (third variable, reverse causation, selection effect) is most plausible, and does the actual study design rule any of them out?
If it doesn't, that's the honest version of the finding, and it's usually still interesting, just not the version that licenses advice.