Subjects · Psychology & Behavioral Sciences

Psychology & Behavioral Sciences: How AI can best tutor this subject

Use AI where it genuinely helps in psychology, which is mostly as a replication-checking and assumption-surfacing instrument, and stop it from teaching you the field's greatest hits as settled fact when a good fraction of them no longer are.

What you'll be able to do: Use AI where it genuinely helps in psychology, which is mostly as a replication-checking and assumption-surfacing instrument, and stop it from teaching you the field's greatest hits as settled fact when a good fraction of them no longer are.

The subject's central problem, and what that implies

Psychology has a danger the other subjects in this corpus don't share to the same degree. A large number of its most confidently taught findings have replication problems: failed to replicate outright, replicated at a fraction of the claimed size, or turned out to rest on a method that doesn't support the conclusion drawn from it. Textbooks have been slow to catch up. Training corpora, built substantially from those textbooks and from decades of confident secondary coverage, inherit the lag and add fluency on top of it.

The result is a specific failure mode: ask an AI tutor to explain a landmark study and it will do so accurately as a summary of what has long been said about the study (correct name, correct year, correct headline finding) while saying nothing about whether the finding held up, because "did this replicate" and "what does this textbook chapter say" are different questions with the same fluent answer if you don't ask the first one separately.

That gives the subject its one reliably productive question, and it's worth learning to ask it about any named finding, on reflex: has this replicated, and what's the current state? Not "is this true": the field rarely gets a clean yes or no, but where the evidence actually sits now, who disputes it, and what changed. Asking this of Milgram's obedience studies, the marshmallow test, ego depletion, power posing, the Stanford Prison Experiment, stereotype threat, or growth mindset gets you a different (and more useful) answer every time than accepting the first fluent summary.

Four further facts about the subject compound this, and each shapes how you should use the tool:

Effect size and statistical significance are different questions, and students are taught to conflate them. "Significant" tells you a result probably isn't noise. It says nothing about whether it's big enough to matter, A huge amount of popularised psychology, including a fair amount of what an AI will say back to you, reports that something was found to be true and never once states how much of a difference it actually made.

The field's early evidence came overwhelmingly from one population: the undergraduates its researchers had cheapest access to. That's not a design choice anyone defended at the time, it's a biographical fact about who does psychology research and where, and it has real consequences downstream. A finding from a 90-minute lab task with American psychology undergraduates is not automatically a finding about people.

The leap from a lab result to a piece of life advice is the field's most common form of overreach, and it happens in a single unmarked sentence: "participants who did X reported more Y" becomes "so you should do X." The study supported the first half. The AI, trained on decades of writing that makes exactly this leap without flagging it, will often make it too, unless told not to.

And psychology's claims are about the reader in a way no other subject on this list is. A theorem doesn't flatter you. A finding about attachment style, self-control, or personality does (or threatens you) and that makes motivated reasoning unusually easy to fall into and unusually hard to notice from the inside. You are the one subject where the AI can be right and you can still not want to hear it, and where an AI trying to be agreeable can hand you exactly the reading you'd have chosen yourself.

What it is genuinely good at here

Checking replication status on demand, for any named study. This is the single highest-value use in the subject and the one almost nobody does by habit. Ask before you learn a finding, not after.

Producing the actual study design, not just the conclusion. Psychology exams and coursework mark the method, sample, design, measures, confounds, ethics, far more than they mark the headline finding. An AI that's made to give you the IV, the DV, the sample and its limits, rather than "researchers found that..", is teaching you the thing you're actually examined on.

Forcing a vague construct into an operational definition. "Self-esteem," "intelligence," "wellbeing," and "aggression" are words before they're measurements, and the measurement is where a study's whole argument either holds or doesn't. Making the definition explicit is cheap and it's where a surprising amount of psychological disagreement actually lives.

Distinguishing correlation from causation with your specific claim, rather than reciting the rule. Everyone can state "correlation isn't causation." Almost nobody can look at a specific finding and say what a third variable, a selection effect, or reverse causation would have to look like to produce it, and that's the skill that's actually tested.

Arguing the reading you didn't want, when the material is about you. If a finding, a quiz result, or a mood pattern concerns your own traits or behaviour, asking for the less flattering interpretation before accepting the first one is a direct countermeasure to the subject's specific motivated- reasoning problem.

Reading a figure properly. A forest plot, an error bar, a confidence interval, an effect-size estimate: these carry the actual evidence, and a conclusion stated in prose without them is a claim you can't evaluate. AI is good at walking you through what a figure is actually showing, once asked.

What it is bad at here, specifically

Stating contested or outdated findings with the same confident register it uses for settled ones. Nothing in the prose style changes between "this is well-established" and "this was influential and later fell apart." You have to ask which one you're getting.

Making the correlation-to-causation leap itself, fluently, in one sentence. Trained on decades of writing that does this, it will often reproduce the habit rather than resist it, unless explicitly told to flag the move every time it happens.

Telling you what you want to hear about yourself. Sycophancy is a general problem with these tools; in psychology it has a subject-specific target. Ask it to interpret your own personality-quiz result, mood log, or attachment pattern and the flattering reading is the statistically likely output unless you ask for the alternative explicitly.

Knowing your specification. Exam boards differ sharply in which evaluation points they reward, some want named methodological critiques (ecological validity, demand characteristics, social desirability bias), others want named alternative studies. Without your syllabus, it will teach a generic version.

Distinguishing "the popularised version" from "what the original paper actually found and claimed". It can do this if asked directly, and it won't volunteer it. The gap between the two is often the entire story, the marshmallow test's predictive power, power posing's hormonal claims, and the Stanford Prison Experiment's "spontaneous" cruelty are all cases where the popular version says considerably more than the underlying study supports.

The shape of a good session

  1. Before learning a named finding: ask for its replication status and current best effect-size estimate against the original claim. Do this before the explanation, not after, otherwise the explanation lands first and does the anchoring.
  2. Before accepting the conclusion: ask for the actual study design. IV, DV, sample, and what was operationalised as what.
  3. For any causal-sounding claim: ask what would have to be true for the correlation to be causal, and whether the design used could actually show that.
  4. For anything about you personally: ask for the less comfortable reading before the comfortable one.
  5. For the evidence itself: ask to see the figure, the effect size, the confidence interval, the forest plot, not just the sentence summarising it.

The instruction to set

For this session you are tutoring me in [topic] in psychology. Rules, for any named study or finding, state its current replication status and what changed since the original if anything did, before explaining it; report effect sizes and confidence intervals, not just whether a result was "significant"; never move from a correlational finding to causal or practical advice without flagging that move explicitly; when the material concerns a trait, mood pattern or quiz result of mine, give me the less flattering interpretation before the flattering one; and tell me plainly when a finding is contested rather than picking a side for me.

That middle clause (flag the causal leap every time) is worth repeating once per topic. It's the single most common thing a fluent answer will do without telling you.

The failure that looks like success

A clear, well-organised, confidently delivered account of a landmark study, which you find easy to remember because it already matches what you expected psychology to say.

You now hold a fact you can recite and cite by name and year, and you have not once been told whether the field still believes it. Under an exam question that asks you to evaluate the study, rather than describe it, you have nothing, because the version you learned never included the part that's actually contested. And the explanation made this worse, not better, because fluency reads as authority.

The tell: you can name the study, the researcher and the finding, and you've never been asked (by the AI or by yourself) whether it replicated.

What it cannot replace

Running an actual study, however small, and hitting the parts a description never contains: recruiting real participants, a consent process, an attrition rate, a manipulation check that quietly fails, a result that doesn't come out clean once you're the one who collected it. That's where "confound" and "demand characteristics" stop being exam vocabulary and start being things that happened to your data.

Use AI around that work, designing it, critiquing the design before you run it, reconciling the messy result afterward. Not instead of running it.

Where to go next

Start with 01 The replication verification drill, which is the subject's one indispensable habit, and 06 Correlation-causation mapping, which is where the field's most common public-facing error actually lives.