Subjects · History & Social Studies
History & Social Studies: Standard-setting comparison
Find the one specific move that separates a top-band history essay from a strong mid-band one, by comparing your work against exemplars side by side rather than against a rubric in the abstract.
What you'll be able to do: Find the one specific move that separates a top-band history essay from a strong mid-band one, by comparing your work against exemplars side by side rather than against a rubric in the abstract.
Why the rubric doesn't tell you
Every mark scheme lists something like "sustained, well-substantiated judgement engaging with the complexity of the issue" for the top band, and something like "developed argument, mostly substantiated" for the one below it. Read both descriptions and you cannot tell what to actually change in your essay. The language scales smoothly; the essays it's grading usually don't.
This is a known limit of rubric-based judging generally: people are unreliable at applying an absolute scale to a single piece of work, and much more reliable at comparing two pieces of work side by side and saying which is better and why. Exam boards use exactly this (comparative judgement) to set real grade boundaries, because graders agree with each other far more when comparing than when scoring against a description. The method here is the same shift, aimed at your own essay: stop reading the rubric, and put your draft next to essays at the grade above and below it.
What that comparison actually turns up, in history specifically
Do this enough times, across enough essays, and one thing shows up disproportionately often: it is very rarely more facts, longer quotations, or more sophisticated vocabulary that separates the top band from a strong mid-band answer. It is nearly always whether the essay names, and properly answers, the strongest evidence or interpretation against its own case.
The mid-band essay makes a good case and stops. The top-band essay makes the same case, then spends a paragraph on the best reason it might be wrong, and explains (specifically, not as a gesture) why the case still stands, or how it has to be qualified. That paragraph is usually short. It is nearly always present in the higher script and nearly always absent in the lower one, and once you've seen this pattern in ten pairs of exemplars, the rubric's "engaging with complexity" stops being vague.
How to run it
- Get exemplars at adjacent bands to the question you're answering, real past-paper exemplars and examiner reports if you have access to them; AI-generated illustrative ones, clearly labelled as constructed and not real scripts, if you don't.
- Read them side by side, not against the rubric text.
- Ask what specifically changes between adjacent bands: not "what's better," which invites vague answers, but the concrete move.
- Apply the finding to your own draft directly.
- Revise for that move, not for general polish.
Here is the essay question: [question]. Write three short illustrative responses to it, mid-band, upper-band, and top-band by [exam board] standards, clearly labelled as constructed examples, not real scripts,
Then tell me specifically what the top-band version does that the mid-band version doesn't. Not "it's more sophisticated": the actual move.
And the honesty clause this needs, because the claim underneath it is narrower than it sounds:
Are these differences based on real mark schemes and examiner reports, or is this your general sense of what "better" essays do? Say which, and where you can, point me to what a real examiner report says about this band gap.
Comparative judgement as an assessment method is well evidenced. Whether an AI-constructed exemplar accurately represents a real grade boundary is a separate, weaker claim, treat generated exemplars as a hypothesis to check against real mark schemes and examiner reports, not as a verified standard.
Worked shape
Question: "How far was the economic crisis the main cause of the French Revolution?"
Mid-band exemplar: makes a sustained, well-evidenced case for the economic crisis (bread prices, state debt, tax structure) and concludes it was the main cause. Competent, focused, nothing wrong with it as far as it goes.
Top-band exemplar: makes the same case, then adds a short paragraph on the political crisis interpretation, the Estates-General, the specific failure of 1789's political institutions, and argues that economic hardship explains grievance but not revolution, since comparable hardship elsewhere in Europe didn't produce one; the political failure explains why grievance became rupture there and then. That paragraph is the entire gap between the two scripts.
Applied to your own draft: find the strongest interpretation you've left unaddressed, write the paragraph, and check, does it actually change your conclusion or qualify it, or is it a sentence that concedes nothing and gets forgotten by the next paragraph?
How this differs from steelmanning (article 03)
Steelmanning is preparation: before you write, you build the strongest opposing case to find out what your position rests on. This method is a finishing check: after a draft exists, you're asking whether the concession actually made it onto the page, and whether it does real work there or sits as a token sentence. Do both, steelman before writing, standard-set compare after, and the second will usually be where you discover the first didn't fully survive into the essay.
Across the subjects
History. As above.
Political science. Evaluative policy essays, where the top-band move is usually engaging with the strongest evidence that the policy actually worked, or actually didn't, against your stated position.
Sociology. Theory-evaluation essays, where the gap is usually whether the strongest critique of your preferred theory gets named and answered.
Anthropology. Essays defending an interpretation of fieldwork, where the move is acknowledging the strongest alternative reading of the same observations.
Geography. "To what extent" essays on causes or impacts, same shape as history.
International relations. Essays assessing a policy or crisis response, where the concession is usually the strongest case that a different choice would also have gone wrong.
Pitfalls
- Trusting generated exemplars as calibrated. They're illustrative. Check against real mark schemes where you can.
- Comparing against the rubric instead of the exemplars. The rubric is what you check the finding against afterwards, not what you compare to directly.
- Chasing every difference at once. Usually one move dominates; find it and drill it before hunting for smaller ones.
- A token concession. A sentence that doesn't change or qualify your conclusion isn't the move, it's decoration wearing the move's clothes.
- The tell: you can point to your concession paragraph and it makes no difference if you delete it. The top-band version's doesn't survive deletion; it's load-bearing.
Try this today
Take an essay question you're preparing. Get two illustrative exemplars at adjacent bands, and ask specifically what the higher one does that the lower one doesn't.
Then find that move missing from your own draft, write the paragraph, and delete it again to check: does your essay's conclusion actually depend on it now?