Lesson 4 - The Two-Hat Technique

Welcome to The Two-Hat Technique

Lesson 3 made self-review honest with better prompting. But there’s a limit to how honest a reviewer can be when it’s grading its own work. The mind that made the draft is attached to it, shares its blind spots, and leans toward defending its choices. The fix is structural: split the work into two roles — a doer that produces, and a reviewer that judges — and keep them genuinely separate. That separation is the two-hat technique, and it’s how you give your check fresh eyes.

By the end of this lesson, you will be able to:

  • Explain why a maker is a biased reviewer of its own work
  • Run the doer and reviewer as separate roles with separate context
  • Give the reviewer only what it needs — the work and the criteria, not the rationale
  • Apply the technique by hand in ChatGPT or Claude today

Let’s start with why grading your own homework goes wrong.


Why a Maker Can’t Fairly Grade Itself

When the same AI (in the same conversation) makes something and then reviews it, two problems stack up. First, it has already committed to its choices — it “decided” the invented book title was fine when it wrote it, so asked moments later whether that title is real, it’s primed to defend, not re-examine. Second, models show a documented self-enhancement bias: they tend to rate their own outputs more favourably than an identical output from someone else. A reviewer that both made the work and is inclined to like its own work is exactly the reviewer you don’t want.

There’s a broader pattern underneath this, which you met last lesson: internal self-correction, with nothing external to push back, is unreliable — a model reasoning in the same context that produced an answer often just reaffirms it. The most reliable checks come from outside the thing being checked. The two-hat technique manufactures that “outside” deliberately: it creates a reviewer that didn’t make the work and doesn’t know how it was made.


Two Hats, Two Contexts

The technique is simple to state: the doer produces the work; a separate reviewer judges it against the criteria; the reviewer never sees the doer’s reasoning. In practice, “separate” means a fresh context — a new chat, or a clearly reset role — so the reviewer starts cold, with no memory of the justifications the doer told itself along the way.

A diagram titled 'Two hats: the maker proposes, a fresh reviewer disposes'. On the left, a Doer (purple) writes the draft and is attached to its own work; it produces two things — 'the Draft, the actual output' and a red box 'looks great to me! its self-praise / rationale'. A dashed vertical line labelled 'fresh context' divides the diagram. A green arrow shows the draft crossing the divide; a red dashed arrow shows the self-praise being blocked with a red cross. On the right, a Reviewer (orange) is given only the draft plus a 'the Criteria, the definition of done' card, and returns a verdict — names the event met, 312 words over 250 not met, invented a title not met. A caption says give the reviewer only the work and the criteria, never the maker's reasoning, so it can't be talked into agreeing.
The two hats. The draft and the criteria cross to the reviewer; the doer's self-praise and reasoning do not. Starting cold, with only the work and the standard, the reviewer catches what the maker was primed to defend.

Notice what crosses the divide and what doesn’t. The draft crosses — the reviewer obviously needs the work. The criteria cross — the reviewer judges against your definition of done. But the doer’s reasoning and self-assessment do not. If you hand the reviewer “here’s the draft, and here’s why I think it’s great,” you’ve just re-infected it with the maker’s bias. Give it the work and the standard, and let it form its own verdict.


How to Run It by Hand

You don’t need any special tooling — you can do this in a browser today, three ways from lightest to strongest:

  • Fresh chat, same model. Finish the draft in one conversation. Open a new chat, paste only the draft and the criteria, and ask for the per-criterion verdict from Lesson 3. The clean context alone removes most of the “defending my own choices” effect.
  • A different model. Have Claude do the work and ChatGPT review it, or vice versa. A different model is even less attached to the draft and less likely to share the exact blind spot that produced an error.
  • An assigned reviewer role. In one chat, explicitly switch hats: “You are now a strict reviewer who did not write this. You care only about whether it meets the criteria. Be skeptical.” Weaker than a fresh context, but better than nothing when a new chat is inconvenient.

The strongest, cheapest habit for important work is the fresh-chat review: it costs one copy-paste and removes the largest source of self-review bias. Reach for a different model when the stakes are higher or when one tool keeps missing the same kind of problem.

One practical warning: the two-hat split only works if you actually keep the contexts apart. If you paste the whole conversation — draft, your prompts, and the doer’s running commentary about how well it’s doing — into the “review,” you’ve handed the reviewer the maker’s mindset and undone the separation. The discipline is to pass only the finished draft and the criteria, stripped of the back-and-forth that produced them. Think of it as handing a colleague the document and the checklist, not the document plus a note that says “I already checked it and it’s great.” The blank slate is the whole point; protect it by being deliberate about what you copy across.

Doer and reviewer are roles, not necessarily two AIs

The two hats can be two AIs, an AI and you, or the same AI in two separate passes. What matters isn’t who wears each hat — it’s that the reviewing hat judges against the criteria without being anchored to how the work was made. Even you review your own writing better after a night’s sleep, for exactly the same reason: the reviewer benefits from distance.


Harborlight: Doer and Reviewer

The owner drafts the newsletter with Claude, going a few rounds until Claude says it’s ready. Instead of trusting that, they open a fresh Claude chat and paste only two things: the draft, and the newsletter’s definition of done. The prompt is the honest-review prompt from Lesson 3 — per-criterion verdicts, evidence required, gaps welcome.

The fresh reviewer, with no attachment to the draft, immediately flags two things the “ready!” doer had waved through: the draft is 312 words (the doer had stopped counting), and it mentions a book — Tidewater — that isn’t in the stock list (the doer had invented it and then trusted its own work). Those go back as corrections; the doer fixes them; a second fresh review passes cleanly. The doer alone would have shipped both errors, glowing with confidence. The second hat caught them for the price of a copy-paste.


Practice Exercises

Exercise 1: What crosses the divide?

You’re setting up a two-hat review. Which of these should you give the reviewer: (a) the draft, (b) the definition of done, (c) the doer’s explanation of why it’s good, (d) the original vague request?

Hint

Give it (a) the draft and (b) the definition of done — the work and the standard. Withhold (c), the doer’s self-justification, which just re-imports the maker’s bias. (d) the vague request is unnecessary and can even mislead; the reviewer should judge against the concrete criteria, not the fuzzy original ask.

Exercise 2: Diagnose the weak setup

Someone says “I asked the AI to write it, then in the same chat asked if it was good, and it said yes.” Which two problems from this lesson are in play?

Hint

Same-context self-review (it’s defending choices it already committed to) plus self-enhancement bias (it favours its own work). Both are fixed by moving the review to a fresh context with only the draft and criteria.

Exercise 3: Pick the separation

For a high-stakes press statement, would you use a fresh chat with the same model, or a different model entirely, as the reviewer? Justify it.

Hint

For high stakes, a different model — it’s the least attached to the draft and least likely to share the exact blind spot that produced any error, giving you the most independent second opinion. A fresh chat with the same model is fine for everyday work; press statements earn the stronger separation.


Summary

Even a well-prompted self-review has a structural weakness: the mind that made the work is biased toward it. A maker is primed to defend the choices it already committed to, and models show a self-enhancement bias — they rate their own outputs more kindly than identical work from elsewhere. The two-hat technique fixes this by splitting the loop into a doer that produces and a reviewer that judges, kept in separate contexts. The draft and the criteria cross to the reviewer; the doer’s reasoning and self-praise do not, so the reviewer starts cold and forms an independent verdict. You can run it by hand today — a fresh chat (cheapest and highly effective), a different model (strongest separation), or an assigned skeptical reviewer role (a lighter fallback). The reviewing hat catches exactly the confident errors the doer waves through, for about the price of a copy-paste.

Key Concepts

  • Two-hat technique — separating the doer (produces) from the reviewer (judges) into distinct roles/contexts.
  • Self-enhancement bias — a model’s tendency to rate its own output more favourably than others'.
  • Fresh context — a new chat or reset role, so the reviewer isn’t anchored to how the work was made.
  • What crosses the divide — the draft and the criteria, never the maker’s reasoning.

Why This Matters

The most confident errors — the invented fact, the missed limit — are exactly the ones a maker rubber-stamps, because it already decided they were fine. A separate reviewer is the cheapest reliable way to catch them, and it scales from a copy-paste in a browser up to two different models double-checking each other. This doer/reviewer split is also the backbone of how automated loops verify themselves, which you’ll see when agents review their own work in Module 5. But even two separate hats can both be fooled the same way — which brings us to the most important lesson in the course: what happens when the check itself is wrong, and how to catch a false “done”.


Continue Building Your Skills

You can now give a check genuinely fresh eyes by splitting the doer from the reviewer — a copy-paste into a new chat removes most self-review bias, and a different model removes even more. Make the fresh-chat review a reflex for anything that matters. Next comes the lesson the whole course has been building toward: when the check itself lies, passing work that isn’t actually done — and how to catch that false “done” before it reaches you.

Sponsor

Keep DATATWEETS free. Help fund practical data, AI, and engineering lessons for learners worldwide.

Buy Me a Coffee at ko-fi.com