Lesson 3 - Generate and Test

Welcome to Generate and Test

Draft-critique-revise improves one attempt, step by step. But for some tasks, the first attempt is a coin flip — a tagline, a creative concept, a name — where two tries can differ wildly in quality and polishing the first one just refines a lucky, or unlucky, start. The pattern for these is generate and test: produce several independent candidates, then test them against your check and pick the best. Instead of improving one attempt serially, you explore several in parallel and let the check choose. It’s a different shape for a different kind of task, and knowing when to reach for it is the skill.

By the end of this lesson, you will be able to:

  • Describe the generate-and-test pattern
  • Recognize high-variance tasks where it beats draft-critique-revise
  • Use your check as the “test” that picks the winner
  • Combine generate-and-test with draft-critique-revise

Let’s look at the pattern.


Explore in Parallel, Then Pick

Generate-and-test has two moves. Generate: from one goal, produce several independent candidates — not variations of one draft, but genuinely different attempts. Test: score each against your check and pick the winner. That’s it. Where draft-critique-revise walks one attempt toward done, generate-and-test lays out several and keeps the best.

A diagram titled 'Generate and test: make several candidates, let the check pick the best'. A blue 'Goal — one prompt' box fans out ('generate N diverse') to three purple boxes 'Candidate A', 'Candidate B', 'Candidate C', each 'an independent attempt'. A 'Test — score each against your check' step scores them: Candidate A 6/10, Candidate B 9/10 with a check mark (highlighted green), Candidate C 5/10. An arrow leads from B to a green 'Winner: Candidate B — then polish it (draft-critique-revise)' box. A caption notes to use it when the first attempt is a coin flip (creative, naming, brainstorming) because diverse candidates beat polishing one, and that serial draft-critique-revise improves one attempt while parallel generate-and-test explores many then picks, and you often combine both.
Generate and test. Produce several independent, diverse candidates from one goal, score each against your check, and keep the winner — then optionally polish it with draft-critique-revise. Parallel exploration, not serial refinement.

The test is your check, doing the same job as always but in a new role: instead of deciding “is this one done?”, it decides “which of these is best?” Everything about building good checks (Module 3) applies — a vague test picks a winner at random, a concrete test picks a genuinely better one. And notice the check is now comparing candidates, which is often easier than judging one in isolation: you may struggle to say whether a single tagline is “good,” but you can usually tell which of three is best. Comparison is a friendlier job for a check than absolute judgment.


When Parallel Beats Serial

The key question is variance: how much do independent attempts differ in quality? That tells you which pattern fits.

Low variance → refine one (draft-critique-revise). For a factual bio or a tidy product description, any reasonable first attempt is close, and the work is polishing it. Generating five wouldn’t help — they’d all be similar, and you’d still have to refine the one you pick. Serial refinement is right.

High variance → generate several (generate-and-test). For a tagline, a campaign concept, a shop name, a creative headline, attempts differ wildly and the first one is a gamble. Polishing one traps you in that idea; the good version might be a completely different concept you’d never reach by revising. Generating several independent, diverse candidates and picking the best explores the space, so you’re choosing the best idea, not perfecting the first one that showed up.

The tell that you should switch to generate-and-test is the feeling from Lesson 1: you start a draft-critique-revise and it feels like polishing a mediocre thing rather than finding the right thing. That’s variance talking — stop refining and generate alternatives instead.

How Many Candidates?

Generating candidates isn’t free — each one costs a little time and, for automated loops, a little money — so “how many” is a real choice, and it follows the budgeting instinct from Module 6. A handful is almost always the right number: three to six diverse candidates give the check a real range to choose from without much waste. Below three, you’re barely exploring; far above six, you’re paying to generate options you won’t seriously consider, and the check has more to wade through. The sweet spot is enough candidates to span the space of reasonable approaches — a few genuinely different angles — not an exhaustive catalog. And the higher the variance and the higher the stakes, the more candidates earn their keep: a throwaway internal name might warrant three, while the tagline going on every poster for a year is worth generating eight and choosing carefully. As with budgets, size the count to what the decision is worth.

Diversity is the whole point

Generate-and-test only works if the candidates are genuinely different. Ask for varied approaches, not slight rewordings: “give me five taglines taking completely different angles — one playful, one elegant, one blunt, one nostalgic, one bold.” Five near-identical taglines give the check nothing meaningful to choose between. The value comes from spanning the space of possibilities, so the winner is the best of a real range, not the least-bad of five similar tries.


Combine the Two

The patterns aren’t rivals — the strongest approach often uses both. Generate-and-test to find the right idea, then draft-critique-revise to perfect it. Generate five diverse taglines, test them, and pick the strongest concept — that’s the exploration. Then take that winner and run a draft-critique-revise on it — tighten the wording, fix the length, match the brand voice — that’s the refinement. You get the best idea from the parallel step and the best execution from the serial step.

This combination is a good default for any creative-but-standards-bound task: explore widely first so you don’t polish the wrong thing, then refine deeply so the right thing is finished well. It’s the same instinct a designer has — sketch many options, choose one, then perfect it — expressed as loop patterns. The order matters: exploring after you’ve committed to refining one idea rarely happens, because sunk effort makes you defend the thing you’ve already polished. Generating the alternatives first, before any polishing, is what keeps you honest about whether your starting idea was actually the best one.


A Harborlight Example

The owner needs a name for a new monthly event. They start a draft-critique-revise on their first idea, “Book Chat,” and it feels like sanding down something dull — the tell that variance is high and they’re on the wrong pattern. So they switch to generate-and-test: “Give me eight event-name ideas in genuinely different styles — cozy, clever, literary, playful, community-focused.” Eight arrive, spanning a real range. They test against a quick check — memorable, not already used locally, fits our warm voice, works on a poster — and one clearly wins: “Chapter & Chat.” Comparison made the choice easy in a way judging “Book Chat” alone never could.

Then they combine: they run a short draft-critique-revise on the winner — checking it’s not taken, drafting a one-line description, confirming it fits the poster. Exploration found the right name; refinement finished it. Had they only refined “Book Chat,” they’d have ended with a polished dull name instead of the good one that was three ideas away.


Practice Exercises

Exercise 1: Which pattern?

For each, is it draft-critique-revise or generate-and-test: (a) fixing the grammar in a finished report, (b) coming up with a slogan for a sale, (c) tightening a bio to under 80 words, (d) naming a new product?

Hint

(a) and (c) are low-variance refinement — draft-critique-revise. (b) and (d) are high-variance creative choices — generate-and-test: produce several diverse options and pick the best, because the good one might be a totally different idea than your first.

Exercise 2: The comparison advantage

Why is “which of these three is best?” often an easier check than “is this one good?”

Hint

Absolute judgment (“is this good?”) has no reference point, so it’s vague and easy to fudge. Comparison gives the check concrete alternatives to weigh against each other, which is a more answerable question — you can usually rank three options even when you can’t score one in isolation.

Exercise 3: Combine them

You’ve generated six campaign concepts and picked the best. What’s the natural next step, and which pattern is it?

Hint

Refine the winner with draft-critique-revise — tighten the wording, fix the length, match the brand voice. Generate-and-test found the best idea; draft-critique-revise perfects its execution. Explore widely, then refine deeply.


Summary

Generate-and-test produces several independent candidates from one goal and picks the best against your check — parallel exploration rather than the serial refinement of draft-critique-revise. The test is your check in a new role, comparing candidates rather than judging one, which is often easier than absolute judgment. The pattern to use depends on variance: low-variance tasks (a factual bio, a tidy description) call for refining one, while high-variance tasks (a tagline, a name, a creative concept) call for generating several — because polishing one traps you in that idea when the good version might be a completely different one. Diversity is essential: candidates must genuinely differ, spanning the space, or the check has nothing meaningful to choose between. And the two patterns combine: generate-and-test to find the right idea, then draft-critique-revise to perfect it — explore widely, then refine deeply.

Key Concepts

  • Generate-and-test — produce several independent candidates and pick the best against your check.
  • Variance — how much independent attempts differ; high variance favors generate-and-test.
  • Comparison vs. absolute judgment — picking the best of several is often an easier check than scoring one.
  • Combining patterns — generate-and-test to find the idea, draft-critique-revise to finish it.

Why This Matters

Knowing generate-and-test frees you from the trap of polishing the first idea that appeared, which is where a lot of mediocre creative work comes from — a decent-but-not-great concept, sanded smooth. For anything with real creative variance, exploring several options and choosing the best simply produces better outcomes, and it’s cheap to do. Combined with draft-critique-revise, it mirrors how good creative work actually happens: many options, one choice, then careful finishing. You’ve now got three patterns — refine one, decompose a big one, explore several. The last pattern lesson scales up further: running loops in parallel and handing work between more than one agent.


Continue Building Your Skills

You can now reach for generate-and-test on high-variance tasks — several diverse candidates, tested against your check, best one wins — and combine it with draft-critique-revise to explore then refine. Watch for the “polishing something mediocre” feeling; it’s your cue to generate alternatives instead. Next, you’ll scale patterns up to more than one loop at once: parallel loops and handoffs between agents, for work that’s too big or too varied for a single pass.

Sponsor

Keep DATATWEETS free. Help fund practical data, AI, and engineering lessons for learners worldwide.

Buy Me a Coffee at ko-fi.com