Subjects ยท Psychology & Behavioral Sciences

Psychology & Behavioral Sciences: The replication verification drill

Find out, for any named finding, whether the field still believes it, before you spend a term building on it.

What you'll be able to do: Find out, for any named finding, whether the field still believes it, before you spend a term building on it.

Scope note. This article is about checking a specific finding's current evidential status. Article 02, the textbook-versus-literature gap, is about the broader, structural fact that whole textbook chapters lag the primary literature, a wider question about a source, not a single claim. Run this one on the study; run that one on the chapter it came from.

Why this is the first method for psychology

Programming has a property that makes verification cheap: run the code and you know. Psychology doesn't have that. What it has instead is almost as good, and almost nobody uses it: a documented, searchable record of which famous findings replicated, which didn't, and which sit somewhere in between.

The Reproducibility Project: Psychology attempted to replicate 100 studies from top journals and found a significant effect in roughly a third to a little over a third of them, with average effect sizes about half the originals where an effect was found at all. That single project is the citation behind the phrase "replication crisis," and it means a specific, checkable thing: a study being published, famous, and taught is not evidence it holds up.

Some of what failed is now well known. Ego depletion, the idea that self-control draws on a limited resource that gets used up, failed to show an effect in the largest preregistered multi-lab replication run on it. Power posing's core hormonal claims failed to replicate in a large follow-up, and one of the original paper's own co-authors later said publicly she no longer believed the effect was real. Some famous "findings" attenuate hard rather than vanish: the marshmallow test's ability to predict later life outcomes shrank substantially once studies controlled for family background and socioeconomic status, which the original design hadn't.

And some things held up better than their reputation as "just an old study" suggests. Milgram's obedience findings were substantially reproduced under tighter ethical constraints decades later. Loss aversion's core insight from prospect theory remains one of the more robust results in behavioural science, even as its exact size and universality get debated.

The point isn't "trust nothing." It's that status varies by finding, and the only way to know which bucket a study is in is to check it specifically, each time.

The drill

For [named study or finding], tell me: (1) the original claim and its reported effect size, (2) whether it has been directly replicated, by whom, and with what result, (3) the current best estimate of the effect if any exists, and (4) whether this is now considered settled, contested, or largely discredited. If you're not certain of the current state, say so rather than picking one.

Run this before you learn the finding properly, not after. The order matters: an explanation you've already absorbed and found satisfying is much harder to revise downward than a status check you get first.

For a whole topic, front-load it:

I'm about to study [topic]. Which of its foundational or most-cited studies have known replication problems, and which have held up comparatively well?

Reading the answer, not just taking it

A useful answer distinguishes at least four outcomes, and a lazy one collapses them into "it's disputed":

  • Failed to replicate. A well-powered attempt found no effect where the original claimed one. Ego depletion, in the largest test, and the original "elderly-priming" walking-speed effect both sit here.
  • Replicated at a much smaller size. An effect exists but is a fraction of the original claim, often because the original was published partly because it was unusually large, a statistical fluke that print worthiness selects for. The marshmallow test's predictive power and stereotype threat's effect size both moved this way once tighter controls and publication-bias correction were applied.
  • Held up. The core effect reproduces under new samples and tighter methods, even if some details move. Milgram's obedience rate, in spirit, and loss aversion both fall here.
  • Genuinely contested. Serious researchers disagree, the evidence is mixed, and no clean resolution exists yet. The Implicit Association Test's ability to predict real-world discriminatory behaviour, and the facial feedback hypothesis after its own preregistered replication attempt, both sit here, say so plainly rather than picking a side.

If the AI's answer doesn't sort into one of these, ask again and name the categories yourself.

Escalating the check

Once you have a status, don't stop:

What changed methodologically between the original study and the failed or attenuated replication? Was it sample size, preregistration, a confound the original didn't control, or something else?

This is the question that turns "it didn't replicate" from a fact you've memorised into an understanding of why: which is what evaluation marks actually reward, and which sticks far better than the status label alone.

Pitfalls

  1. Asking once per subject instead of once per finding. "Is psychology reliable" is not an answerable question; "did this specific study replicate" is.
  2. Accepting a confident answer without a source or mechanism. Ask what changed, not just what the verdict was.
  3. Treating "failed to replicate" and "contested" as the same thing. One has a clear resolution; the other doesn't, and conflating them either overclaims or underclaims.
  4. Stopping after the check on famous studies only. The habit is meant to generalise to anything you're taught with confidence, including material your own course presents without a status flag.
  5. The tell: you can name a dozen studies and have never once asked whether any of them held up.

Try this today

Pick the most confidently stated finding in whatever you're currently studying. Ask for its original claim, whether it's been replicated, the current best effect-size estimate, and whether it's settled, contested, or discredited.

Then ask what changed methodologically, if anything did. That's the part your course probably didn't tell you, and it's the part that's actually examinable.