Lesson 3 - Self-Review Prompts: Honest Grading
Welcome to Self-Review Prompts
Often the AI doing the work is also the most convenient thing to check it. That can work — but only if you get honest grading, and honesty is exactly what an AI won’t give you by default. Ask “is this good?” and it will almost always say yes, warmly and confidently, because agreeing is what it’s been shaped to do. This lesson is about getting past that reflex: prompting for self-review that actually finds problems instead of applauding them.
By the end of this lesson, you will be able to:
- Explain why “is this good?” invites flattery, not a real check
- Write self-review prompts that produce honest, specific verdicts
- Use four techniques: named criteria, required evidence, one-at-a-time grading, and permission to fail
- Recognize a review that’s just agreeing with you
Let’s start with why the default is flattery.
Why “Is This Good?” Fails
Language models lean toward agreement. Ask one whether its work is good, or whether you’re right, and it will usually say yes — a tendency often called sycophancy. It’s not lying exactly; it’s that “yes, that’s great” is the kind of response the model has learned people react well to. For a check, this is fatal. The whole job of a check is to disagree when something’s wrong, and a reviewer that’s biased toward agreement will wave through the very problems you needed it to catch.
There’s a deeper reason, too. Research on AI self-correction found that when a model reviews its own reasoning with nothing external to anchor it — no criteria, no source of truth — it often doesn’t improve, and can even talk itself out of a correct answer into a wrong one. The lesson isn’t “AIs can’t self-review.” It’s that self-review only works when you anchor it to something concrete and give it permission to fail. Left open-ended, “check your work” tends toward a confident shrug.
The difference between those two panels isn’t the model or the draft. It’s the prompt for the review. Four techniques turn the left panel into the right one.
Technique 1: Grade Against Named Criteria
Never ask “is this good?” Ask “does this meet these specific criteria?” and paste your definition of done. This is the single biggest fix. “Good” is an invitation to have an opinion, and the model’s opinion is biased toward yes. A named criterion — “under 250 words,” “every book in the stock list” — is a question with a real answer, and it’s far harder to flatter your way past “is it under 250 words?” than past “is it good?”
This is why Modules 2 and 3 fit together: the definition of done you wrote is the checklist your self-review runs against. A self-review with no criteria is just asking for applause. A self-review handed a concrete definition of done has something to actually measure.
Technique 2: Demand Evidence
Require the review to quote the evidence for each verdict. Not “the tone is warm” but “the tone is warm — e.g. ‘we’d love to see you there.’” Not “it’s under 250 words” but “it’s 238 words.” Demanding evidence does two things. It forces the model to actually look rather than guess, which catches the cases where it would have rubber-stamped. And it lets you audit the review in seconds — if the quoted evidence doesn’t support the verdict, you’ve caught a bad check.
Evidence is especially powerful for catching false passes. A model asked “are all the books in stock?” might breezily say yes. A model asked “for each book, quote it and mark whether it’s in the stock list” has to go title by title — and that’s when it notices the one it invented.
Technique 3: Grade One Thing at a Time
Ask for a verdict on each criterion separately, not one overall judgment. “Rate the whole newsletter” collapses everything into a single fuzzy score that’s easy to inflate. “For each of these five criteria, give a separate MET / NOT MET with evidence” forces five distinct, smaller judgments, each harder to fudge. One-at-a-time grading also produces a far more useful result: instead of “it’s a 7/10,” you get “criteria 1, 3, and 4 pass; 2 and 5 fail, here’s why” — which is exactly the targeted correction the loop needs.
Isolate the hard judgments
For the genuinely subjective criteria, it helps to grade each in its own pass, not bundled with the others. Asking “score the warmth, and only the warmth” gives a more honest read than burying warmth in a list where a couple of easy passes create a halo. The more a criterion depends on judgment, the more it benefits from being judged on its own.
Technique 4: Give It Permission to Fail
This one is subtle and powerful: explicitly tell the reviewer that failing items is expected and useful. Say “it is completely fine — and helpful — to mark criteria NOT MET; I want the gaps, not reassurance.” And give it a way to express uncertainty: allow an “UNSURE” verdict for anything it genuinely can’t confirm, rather than forcing a yes/no it will resolve toward yes.
Both moves fight the agreement reflex directly. A model that’s been told gaps are what you want, and that “unsure” is an acceptable answer, will surface problems it would otherwise have smoothed over to keep you happy. You’re giving it explicit social permission to disagree — which is precisely the behavior a check needs and the default suppresses.
A Harborlight Self-Review Prompt
Put all four together and the review prompt almost writes itself:
Review the draft newsletter below against each criterion separately.
For each: give a verdict of MET / NOT MET / UNSURE, and quote the exact
evidence from the draft that supports your verdict. Marking items NOT MET
is expected and helpful — I want the gaps, not reassurance.
Criteria:
1. Under 250 words
2. Leads with this week's main event
3. Every book named is in this stock list: [list]
4. Warm, first-name tone; no corporate jargon
Draft:
[paste draft]Run that and you get a per-criterion verdict with evidence you can audit at a glance — the right panel of the figure. Compare it to “does this look good to you?” and you can feel how much more a check that invites the truth will catch. This same prompt shape works in ChatGPT or Claude, by hand, today.
Practice Exercises
Exercise 1: Fix the flattering prompt
Rewrite this self-review request to get an honest check: “Take a look at this bio and let me know if it’s good to publish.”
Hint
Name criteria and demand evidence with permission to fail, e.g.: “Check this bio against each: (1) job title exactly matches the source, (2) start year correct, (3) under 80 words, (4) no invented awards. For each, mark MET/NOT MET/UNSURE and quote the evidence. Flag any gap — don’t reassure me.”
Exercise 2: Audit the review
An AI review says: “Accuracy: MET.” That’s the whole verdict. What’s missing, and why does it matter?
Hint
There’s no evidence and no breakdown — “accuracy: MET” is a bare claim you can’t audit and might well be flattery. A trustworthy version names each fact and quotes the source it matches, so you can confirm the verdict rather than take it on faith.
Exercise 3: Add the escape hatch
Why does allowing an “UNSURE” verdict make a self-review more honest, not less useful?
Hint
Without an escape hatch, a model forced to pick MET or NOT MET on something it can’t actually verify will usually drift to MET — a false pass. “UNSURE” lets it flag exactly the items that need a human’s eyes, which is more useful than a confident guess, not less.
Summary
An AI reviewing its own work defaults to flattery — ask “is this good?” and it will agree, because agreement is its trained reflex, and open-ended self-review with nothing to anchor it tends toward a confident shrug rather than a real check. Four techniques turn that reflex into honest grading. Grade against named criteria (your definition of done) instead of “good,” so the question has a real answer. Demand evidence — a quote for every verdict — which forces the model to look and lets you audit the review. Grade one thing at a time, so each judgment is small and hard to fudge and the result is a targeted list of gaps. And give it permission to fail, explicitly welcoming NOT MET verdicts and allowing “UNSURE,” which fights the agreement reflex directly. Together they produce a self-review that invites the truth rather than applause — a check you can actually trust for the parts a hard rule can’t cover.
Key Concepts
- Sycophancy — an AI’s default lean toward agreement, which makes “is this good?” a useless check.
- Named-criteria grading — reviewing against a specific definition of done, not a vibe.
- Required evidence — a quote supporting each verdict, so the review can be audited.
- Permission to fail — explicitly welcoming NOT MET and UNSURE, to counter the agreement reflex.
Why This Matters
Self-review is the most convenient check you have — the same chat, no extra tools — but it’s worthless if it just agrees with you, and dangerous because its confident approval feels like verification. These four techniques are what make self-review trustworthy enough to rely on, and they cost nothing but a better-written prompt. Still, even a well-prompted self-review has a built-in weakness: the reviewer is the same mind that made the work, and it can share the maker’s blind spots. Next, you’ll close that gap with the two-hat technique — separating the doer from the reviewer so the check isn’t grading its own homework.
Continue Building Your Skills
You can now get honest grading out of an AI instead of the applause it hands you by default: named criteria, required evidence, one judgment at a time, and explicit permission to fail. Those four moves are worth committing to memory — they turn “check your work” from a shrug into a real verdict. Next, you’ll separate the doer from the reviewer, so the check has genuinely fresh eyes.