Subjects · Psychology & Behavioral Sciences
Psychology & Behavioral Sciences: Forced operationalisation
State exactly what a study measured when it claimed to measure "self-esteem" or "intelligence", and see how much of the disagreement in the subject is actually a disagreement about that.
What you'll be able to do: State exactly what a study measured when it claimed to measure "self-esteem" or "intelligence", and see how much of the disagreement in the subject is actually a disagreement about that.
The word that isn't the measurement
Every introductory psychology term names something you can't observe directly. Nobody has ever measured "aggression," "self-esteem," "wellbeing," "intelligence," or "attachment security." What gets measured is always a stand-in: a score on a questionnaire, a count of a specific behaviour in a specific setting, a response time, a hormone level. The stand-in is the operational definition, and it is doing all the work the abstract word can't.
This matters more in psychology than in most subjects because the gap between the word and the measurement is unusually wide, and unusually easy to forget about once a finding is stated in prose. "Aggression" in one study might mean the intensity of a noise blast a participant chooses to deliver to a fictional opponent in a lab game. In another it might mean parent-reported playground behaviour. These are not the same thing wearing the same label, and a huge fraction of apparent disagreement between studies on "the same topic" turns out to be disagreement about which one they measured.
A claim can be perfectly well defined: everyone agrees roughly what "self-esteem" means in conversation, and still be poorly operationalised, because agreeing on the concept doesn't mean agreeing on how to turn it into a number. This is a different failure from vagueness (Definition forcing, definition forcing): you can know exactly what you mean and still not have said how you'd measure it.
The constructs worth interrogating on sight
Intelligence. Usually operationalised as an IQ test score. What that score does and doesn't predict, and what it leaves out, creativity, practical skill, several cultures' own concepts of competence, is most of the field's internal argument about the term.
Self-esteem. Usually a self-report questionnaire (agreeing with statements like "I feel I have a number of good qualities"). This makes it vulnerable to social desirability bias and to the difference between how someone rates themselves and how they actually behave.
Aggression. Ranges from a lab proxy (noise blast intensity, willingness to give an opponent a larger "punishment") to real-world measures (arrest records, peer or teacher reports). The lab proxies are what most classic findings used, and they generalise to real aggression only as well as the proxy resembles it, which is exactly the kind of assumption article 07 covers.
Wellbeing. Sometimes a single life-satisfaction rating, sometimes a multi-item scale covering mood, purpose and relationships. Studies claiming to measure "the same thing" frequently used different scales measuring overlapping but distinct constructs.
Attachment security. Originally operationalised through a specific laboratory procedure, a child's behaviour during brief separations from and reunions with a caregiver, in an unfamiliar room, later generalised into adult self-report questionnaires about relationship style. The adult questionnaire version and the original toddler-lab-procedure version measure related but not identical things, and a lot of pop-psychology writing on "attachment styles" quietly uses the adult self-report version's looseness while borrowing the original's scientific-sounding authority.
The drill
The claim is: "[claim, e.g. 'people with higher self-esteem are less aggressive']." How was each construct actually operationalised in the study or studies behind this, what specific measurement, scale, or observed behaviour stood in for the abstract term? What would a different, equally reasonable operationalisation of the same construct have measured instead, and would the finding likely still hold?
That last question is the one that does the work. If a finding about "aggression" only shows up with the noise-blast proxy and not with teacher-reported behaviour, you've learned something real about the finding's limits, not that it's false, but that it's narrower than the one-word label suggested.
Turning it back on your own coursework
This is where the method pays off fastest, because operationalisation marks are cheap and commonly missed:
Here's my study design: I want to measure [construct]. Propose three different ways to operationalise it, and for each, name one thing it would capture well and one thing it would miss.
Doing this before you run or write up a study, rather than after, is the difference between a design section that says "we measured stress" and one that names a specific validated instrument and defends the choice, which is what coursework marking criteria actually reward.
Pitfalls
- Treating agreement on the word as agreement on the measurement. Everyone nods at "self-esteem." Few people notice they mean different instruments.
- Comparing studies that used different operationalisations as if they measured the same thing. This is where a lot of apparent contradictions in the literature dissolve, or turn out to be real.
- Accepting a lab proxy as equivalent to the real-world construct without asking. The noise-blast measure of aggression is a genuine research tool and also a genuinely narrow one.
- Stopping at "how was it measured" without asking what a different measurement would have found. That's the question that reveals how much the finding depends on the specific choice.
- The tell: you can define a construct in a sentence and can't say how any specific study actually measured it.
Try this today
Take a claim from your current topic that uses an abstract construct, self-esteem, intelligence, aggression, stress, wellbeing. Ask how it was actually operationalised in the study behind the claim, and what a different, equally defensible operationalisation would have measured instead.
If the answer changes the finding's plausibility, you've found where the argument actually lives.