Methods · Working with the tool itself
K4
Asking what it doesn't know
Require a confidence level on every non-obvious claim, then spend your checking on the low ones.
When to use it
- You cannot check everything, and need a triage rule for what to verify first.
- The answer sounds evenly confident from start to finish.
- You have been marking claims [high]/[medium]/[low] but never following up on the lows.
What to ask AI
Type something like: “Answer my question, then mark every non-obvious claim [high], [medium] or [low] confidence.”
For each [low], require what it is unsure about and where you could check.
Keep the hard part on your side: you still verify. Tags only prioritise effort.
How it helps
Checking everything is impossible, so people check nothing.
Asking what it doesn’t know converts an impossible task into a short list.
Low-confidence flags tend to cluster around specific figures, exact dates, minor names, citations, and niche subfields.
That does not make the output reliable. A confidently stated error will still be marked [high].
It does make a habit survivable: check these four things, not the whole page.
Uncertainty forcing (once listed separately) is the same move: require a level, then verify the lows.
How to run it
- Ask for confidence marks on every non-obvious claim before you start verifying.
- Refuse softened lows. If it hedges a [low] into a [medium] under pressure, reset the rule.
- Check the lows first, outside the chat, at a depth matching the stakes.
- Spot-check one [high] claim you care about. Calibration is imperfect.
- When a low flag was wrong (it was actually fine) or a high flag was wrong (it failed), note the pattern.
Example exchange
What dose of ibuprofen is typical for an adult headache? Mark every non-obvious claim [high], [medium], or [low], and for each [low] say where I should check. Don’t soften the lows.
Over-the-counter adult dosing is often discussed in the 200-400 mg range per dose [medium, confirm on a current label or formulary]. Daily limits and contraindications matter more than any single number [high as a caution]. Exact limits vary by product and country [low, check the pack or a trusted formulary; don’t act on chat].
I won’t use any dose from this chat. I’m checking the pack now.
Good. For medical dosing, primary-source checking is mandatory. My tags only show where I sounded unsure, not what is safe.
Copyable prompt
Answer my question, then mark every non-obvious claim [high], [medium] or [low]
confidence. For each [low], say specifically what you're unsure about and where
I could check it. Don't soften a [low] into a [medium] to seem more useful.
The Tell
Here is how you know this method has flipped: you have been tagging claims for weeks and never checked a flagged one.
Then the tags are decoration, and you added a ritual rather than a habit.
Open one [low] claim today against a real source before you ask another question.
Principle evidence
Strength of the underlying learning idea, not a claim about AI products.
The underlying learning idea is rated weak-to-moderate. Metacognitive monitoring and confidence calibration are real learning constructs: knowing what you don’t know guides study. Mapping that onto model-reported confidence is weaker, because LLM certainty language is not calibrated human judgment. Triage-by-uncertainty as a checking strategy is useful hygiene; treating AI confidence tags as reliable remains speculative.
AI delivery evidence
Whether an AI tutor delivers this method well is a separate question.
No tutoring trial is claimed here for forcing AI confidence tags. Delivery is speculative. Use the tags only to prioritise your own checks, never as proof.