Methods · Assessment & diagnosis · Short guide
H12
Comparative multi-grader
Have one piece marked by harsh, lenient, and rubric-strict graders, then analyse the differences so you learn where grades are stable versus noise.
When to use it
- One mark feels like absolute truth and you spiral or inflate from it.
- You want to separate stable critique from taste or severity.
- You have a short piece ready and can afford three labelled reads.
What to ask AI
Paste work and rubric. Ask for three separate personas: harsh, lenient, rubric-strict, scores and comments kept separate.
Build a consensus table yourself: agreed strengths/weaknesses versus contested.
Ask it only to critique whether you correctly separated stable versus contested feedback.
How it helps
A single mark feels like truth. Multi-grader comparison shows variance.
What every grader agrees on is your real signal. What only one voice cares about may be taste or severity.
You learn to weight stable critique and not spiral over one harsh voice, or inflate from one kind one.
How to run it
- Submit one artifact to all three personas.
- Keep scores and comments labelled separately, no merged voice.
- Build a consensus table: agreed versus contested.
- Revise from agreed items first.
- Note severity patterns for emotional calibration.
Example exchange
Here's my work and the rubric. Mark it three ways, harsh, lenient, and strict-to-rubric, and keep the scores separate.
Harsh: 48. Lenient: 72. Rubric-strict: 61. Comments labelled per persona.
Agreed across all three: weak counter-argument and a missing definition. Contested: tone and “ambition.” I'll revise the agreed items first.
That separation looks right. Don't average the scores, fix the two agreed flaws first.
Copyable prompt
Work: [ ]. Rubric: [ ].
Mark this three times as separate personas: HARSH, LENIENT, RUBRIC-STRICT.
Keep scores and comments separate: no merged voice. Then wait while I build a
consensus table. Critique only whether I correctly separated stable vs contested
feedback.
The Tell
Your mood tracks the last persona you read, not the consensus table.
Averaging scores without reading comments also fails the method.
Obeying only the harsh grader is severity theatre, not diagnosis.
Principle evidence
Strength of the underlying learning idea, not a claim about AI products.
The underlying learning idea is rated weak. Multiple grader personas illustrate rater variance: useful assessment-literacy awareness, thin as a learning method beyond that.
AI delivery evidence
Whether an AI tutor delivers this method well is a separate question.
AI grader personas are speculative role-play. Use them to practise separating stable from contested critique, not to predict your real mark.