Methods · Working with the tool itself · Short guide
K12
Output-diff critique
Compare two AI drafts and say which is better and why: training criteria, not passive acceptance of the latest regenerate.
When to use it
- You have two AI drafts and were about to keep the longer one by default.
- You need to practise evaluative criteria, not passive acceptance of the latest regenerate.
- A coursework draft needs a merge you write yourself.
What to ask AI
Type something like: “Draft A: [ ]. Draft B: [ ]. Do not declare a winner. Ask me to rank and state criteria.”
Then have it attack inconsistent criteria.
Keep the hard part on your side: you produce the merged version.
How it helps
Comparative judgment trains criteria more than reading a single polished draft.
Diffing two AI outputs is that practice on tool text.
The risk is aesthetic preference: longer, denser, or more confident wins without reason.
Force criteria first, then choose, then merge in your own words.
How to run it
- Generate or collect two drafts for the same brief.
- Before reading preferences aloud, write your criteria: accuracy, structure, audience fit, evidence.
- Rank with reasons. Refuse a winner declared by the model.
- Let it probe contradictions in your criteria.
- Merge yourself. Do not ask it to autogenerate the final as a third mystery draft.
Example exchange
I’ve attached Draft A and Draft B. Please don’t declare a winner. Ask me to rank them with clear criteria first.
What matters more for this assignment: mechanistic accuracy, clarity for a peer, or citation density?
Accuracy first. A wins on mechanism; B wins on clarity. I favoured B at first just because it was longer.
Your criteria just caught an aesthetic bias. When you merge, keep A’s mechanism and B’s signposting, and write the merge yourself.
Copyable prompt
Draft A: [ ]. Draft B: [ ].
Do not declare a winner. Ask me to rank and state criteria. Then attack
inconsistent criteria. I will produce a merged version myself.
The Tell
Here is how you know this method has flipped: you always keep the longer draft.
Length became your only criterion.
Write criteria before you look, then merge on purpose.
Principle evidence
Strength of the underlying learning idea, not a claim about AI products.
The underlying learning idea is rated weak-to-moderate. Comparative judgment and peer-assessment-adjacent skills train evaluative criteria more than passive acceptance of a single draft. Diffing two AI outputs is that practice on tool text. AI-output tournaments as controlled pedagogy remain speculative.
AI delivery evidence
Whether an AI tutor delivers this method well is a separate question.
Models can help probe your criteria. They should not silently pick winners. Delivery speculative; your stated criteria and self-written merge are the learning.