Subjects ยท Psychology & Behavioral Sciences

Psychology & Behavioral Sciences: Figure reading walkthrough

Read an effect size, an error bar and a forest plot as evidence, instead of skipping to the sentence that tells you whether the star is there.

What you'll be able to do: Read an effect size, an error bar and a forest plot as evidence, instead of skipping to the sentence that tells you whether the star is there.

The question a p-value doesn't answer

"Was it significant?" and "was it big?" feel like the same question. They aren't, and psychology is the subject where the gap between them does the most damage, because a huge amount of writing about the field, textbooks, journalism, and a fluent AI summary alike, reports only the first.

A p-value under.05 tells you a result probably isn't pure noise, at a chosen threshold, in that sample. It says nothing about how large the effect is, nothing about whether it's large enough to matter for anything you'd actually do, and (with a large enough sample) it will eventually flag effects too small to notice in practice. A study of ten thousand people can report a "significant" difference of a fraction of a percentage point. The word "significant" in a psychology paper is a technical term about noise, not an everyday claim about importance, and the two meanings get swapped constantly.

The only way to answer the question that actually matters, how big, how certain, how much does this move anything, is to read the figure.

The three figures worth being able to read cold

The error bar. Usually a confidence interval around a mean or an effect estimate. A wide bar means the true value could plausibly be almost anywhere across a broad range; a narrow one means the estimate is precise. Two groups whose bars overlap substantially are not clearly different, whatever the prose above the chart claims.

The effect size, most often Cohen's d. A standardised measure of how far apart two group means are, in units of standard deviation. Rough conventions treat 0.2 as small, 0.5 as medium, 0.8 as large, conventions, not laws, and worth treating as approximate. A study can report p <.001 and d = 0.1 in the same sentence: highly unlikely to be noise, and tiny. Both things are true at once, and only the second tells you whether it matters.

The forest plot. A meta-analysis's summary figure: one horizontal line per included study, showing that study's effect estimate and confidence interval, with a diamond at the bottom showing the pooled estimate across all of them. Three things to check every time: whether the individual study lines cluster together or scatter wildly (scatter suggests real disagreement between studies, not just individual noise); how wide the pooled diamond is (a narrow diamond is a confident pooled estimate; a wide one isn't); and whether the whole picture shifts noticeably when small or low-quality studies are excluded, which is often reported as a sensitivity check underneath.

The walkthrough

Here's a figure from a psychology study [describe or paste it]. Walk me through it: what's on each axis, what does each point or bar represent, what does the error bar or confidence interval show, and what would I have to believe to conclude the groups are meaningfully different rather than just possibly different?

Then push past the walkthrough into the size question, which is the one that usually gets skipped:

What's the effect size here, in standard terms (Cohen's d, r, or odds ratio, whichever applies)? Using the rough small/medium/large conventions, where does it fall, and separately, is that a size that would actually matter for anything I'd do differently in real life?

Those are two different questions and the answer to the first doesn't answer the second. A small but reliable effect might matter enormously at a population level (public health) and barely at all for an individual decision; a large effect in an artificial lab task might not transfer to anything outside it.

Reading a forest plot specifically

Here's a forest plot from a meta-analysis on [topic]. Which individual studies show effects in the opposite direction from the pooled estimate? How wide is the confidence interval on the pooled diamond compared to the individual study intervals, and what does that tell me about how confident I should be in the overall number?

This is where a lot of "the research shows" claims get shakier than the headline sentence suggests, a pooled estimate can look decisive while sitting on top of individual studies that disagree substantially with each other and with it.

Where this connects to the rest of the subject

Reading the figure is what makes article 01's replication check concrete rather than verbal. "It replicated, but at a smaller size" is a sentence; seeing the original study's bar next to the replication's, with the second one noticeably shorter and its interval wider, is the same fact as something you can actually evaluate rather than take on trust.

Pitfalls

  1. Reading only the caption or the prose summary. Both routinely say more than the figure supports.
  2. Treating "significant" and "large" as synonyms. They aren't related in the way the words suggest.
  3. Ignoring interval width. A point estimate without its interval is half a fact.
  4. Skipping the forest plot's individual lines. The pooled diamond can hide real disagreement between the studies that produced it.
  5. The tell: you can state a study's conclusion and have never looked at its actual chart.

Try this today

Find a figure from something you're currently studying, a bar chart with error bars, or a reported effect size. Ask what the interval means, what the effect size is in standard terms, and whether that size would actually matter for anything real, separate from whether it was statistically significant.