A six-gate decision framework for separating small, valuable AI projects from fast demos that leave adoption, integration, and ownership for later.
A team has two weeks, a small budget, and permission to test one AI idea. Which proposal deserves the money?
The usual answer is the one that looks easiest to demonstrate. Perhaps a model can summarize a document, classify a request, draft a response, or query a small knowledge base. A polished screen appears quickly, the steering group sees movement, and the project earns the label of an early success.
But ease of demonstration is a weak selection rule. The smallest technical build may depend on a process that nobody owns, data that cannot be used, users who have no reason to change, or review work that costs more than the automation saves. It can be fast and still be strategically poor.
I would qualify a small AI or data project through six gates before funding it. The gates are not meant to turn a modest initiative into a large governance exercise. They are a way to keep “small” attached to the whole change, not only to the code.
Use three labels for each gate:
| Gate | The decision to make | Evidence worth accepting | Warning sign |
|---|---|---|---|
| Outcome | Is there one operational result worth changing? | A baseline for time, cost, quality, risk, or service completion | “People will be more productive” |
| Boundary | Can one end-to-end slice be tested without pretending it is the whole system? | Named users, inputs, outputs, exceptions, and exclusions | A narrow model task inside an unchanged, unmeasured process |
| Readiness | Can the receiving workflow use the result now? | Accessible data, available users, permissions, review capacity, and a process owner | Adoption, integration, or policy decisions deferred until after the demo |
| Evidence | Will the test answer a decision? | A comparison method, representative cases, and a pass/change/stop threshold | Success defined as completing the build |
| Operation | Can someone support the useful version? | An owner, run-cost estimate, monitoring plan, fallback, and support path | “The platform team will handle it later” |
| Exit | Can the organization stop or reverse the test cleanly? | Time limit, spending cap, rollback path, and treatment of data and artifacts | A pilot that quietly creates an indefinite dependency |
Do not add the labels into a total score. A blocked access-control requirement should not be averaged away by strong executive enthusiasm. A missing baseline might be repairable in a week; a missing legal authority to process the data may stop the idea entirely. The point is to expose the condition, decide what it means, and avoid false precision.
The six gates also separate this decision from general project prioritization. First, find the business constraint before building with AI. Then ask whether a particular small intervention can change that constraint under current operating conditions. A valuable problem can still have a badly timed solution.
Project size is often estimated from the visible build: screens, model calls, pipelines, integrations, or engineering days. That leaves out the work required to produce credible learning.
A genuinely small project includes enough of the path to observe:
Imagine an accounts-payable team considering AI-assisted invoice coding. Extracting supplier, amount, and purchase-order fields from ten clean PDFs is a small model demonstration. A useful small project would instead take a bounded sample of representative invoices, validate the extracted fields, route low-confidence cases to a person, place accepted records into a test queue, and measure correction time and completion. It may use fewer document types, but it covers a complete learning loop.
That vertical slice is usually more informative than a broader interface with no operational handoff. It reveals whether the format variation is manageable, whether reviewers understand the confidence signal, whether the integration preserves the right records, and whether time saved in extraction simply becomes time spent checking.
The UK Government Digital Service’s guidance on using performance data to improve a service makes an important operational point: measurement should be designed from the start and should show whether people can complete the task, choose to use the service, and receive something that meets their needs. An AI test that stops at output quality cannot answer those questions.
Teams often call data access, user training, process ownership, security review, and integration “dependencies.” That language makes them sound external to the project. For qualification purposes, they are part of the product.
Consider a support-response assistant. The model may draft good answers, but value still depends on several conditions:
If these conditions are absent, the team has not discovered minor implementation details. It has discovered that the business process is not ready to receive the proposed capability.
Readiness does not require enterprise-wide transformation. A motivated group of ten users, one approved knowledge collection, a manual handoff, and a named supervisor may be sufficient. The standard is proportional: can the selected slice produce honest evidence without hiding essential work?
This is why pull-first project scoping is useful. Define the future workflow that the test actually needs, then pull in only the data, tools, controls, and human decisions required for that workflow. The boundary stays narrow without becoming artificial.
An initiative can have a persuasive benefit estimate and still produce little realized value. The gap is often explained by adoption and operating burden.
A simple way to challenge the estimate is:
Realized benefit = eligible volume × successful use rate × unit benefit − new operating burden
Each term deserves evidence.
Eligible volume is the work the system can actually touch, not the department’s total workload. If the first version handles only English-language requests with complete account data, use that volume.
Successful use rate combines adoption with task completion. A tool used by half the intended group and accepted in half of those cases does not affect every eligible item.
Unit benefit might be minutes of work avoided, errors prevented, faster resolution, or additional capacity. Use a baseline and state whether the time can truly be reassigned. Five minutes technically saved but absorbed by extra checking is not five minutes of capacity created.
New operating burden includes human review, exception handling, support, model or platform consumption, observability, data maintenance, security work, training, and the cost of correcting failures. Internal labor counts even when it does not create an invoice.
Suppose a document assistant saves a reviewer four minutes on accepted cases. That sounds promising. If only 40 percent of documents are eligible, users accept half the suggestions, and each case adds a minute of verification, the portfolio estimate changes quickly. The project may still be worth testing, but now the test has a useful question: can the team improve eligibility and accepted-use rates without increasing error or review burden?
Microsoft’s 2026 guidance on defining value before building an agent similarly starts with a quantitative before-state, distinguishes sponsors, operators, users, and governors, and treats change burden as part of effort. That is a healthier model than calculating return from model accuracy alone.
Many pilots begin with the easiest path because teams want early momentum. The result is evidence about something nobody seriously doubted.
If the main uncertainty is whether a model can summarize a standard report, twenty ideal reports add little knowledge. If production value depends on messy attachments, restricted records, multilingual inputs, or a difficult handoff, the pilot should encounter a controlled version of that condition early.
This does not mean starting with the highest-risk live action. A team should not give an untested agent permission to change customer records just to make the experiment realistic. It can use a read-only integration, a shadow workflow, sandbox records, replayed cases, or human confirmation. The test should preserve safety while confronting the important uncertainty.
AWS Prescriptive Guidance now frames a generative AI proof of concept as validation across business value, data readiness, technical feasibility, and delivery risk. That is more demanding than showing that a model can return a plausible answer, but it also makes the result more useful. A failed test of the critical assumption can save far more than a successful test of an easy one.
Before work starts, complete this sentence:
We are spending this time and money to learn whether ______, because the answer will determine whether we ______.
Examples include:
If the second blank has no real decision, the activity is exploration. Exploration can be worthwhile, but it should not be sold as a business result.
A demonstration can end with a presentation. A useful change enters somebody’s working life.
Name at least four responsibilities before approval, even if one person holds more than one:
Ownership is not a list of people invited to a meeting. Each owner needs a decision they are allowed and expected to make.
For example, the workflow owner can decide that a category of requests must return to manual processing. The technical owner can disable a prompt or tool version after a regression. The outcome owner can stop funding when the value threshold is missed. The risk owner can prevent expansion beyond the approved data boundary.
This structure also protects technical teams from inheriting every consequence of adoption. Engineers can operate a service, but they cannot make users adopt it, correct every source document, define every business exception, or decide which customer outcome is acceptable. Internal systems need enduring product responsibility; treating internal AI systems like products explains why the owner, user group, operating boundary, and feedback loop should survive the launch.
A small project’s approval can fit on one page. Call it a funding contract, experiment brief, or decision record. The name matters less than the commitments.
Record:
Review the contract when evidence arrives, not only when the calendar ends. A project can stop early because the value is too small, change because users reveal a different need, or expand because the result crosses a pre-agreed threshold. Learning is progress when it changes a decision.
The exit conditions are especially important. Without them, an inexpensive trial can become a permanent manual process, a forgotten cloud bill, an unsupported integration, or a vendor dependency that nobody deliberately approved. Strategy includes choosing what not to build, and it also includes stopping small work that no longer earns attention.
Qualification is not a contest between immediate approval and rejection. Each gate can lead to one of four sensible outcomes.
Fund now when the outcome matters, the slice is complete enough to learn, the workflow can receive it, and the risks are bounded.
Repair readiness first when the idea is credible but a short, owned task—such as measuring the baseline, approving a data set, or recruiting users—must happen before the technical test.
Run a learning experiment when the organization wants technical knowledge but cannot yet claim operational value. Keep the scope, budget, and language honest.
Decline or defer when the project depends on unresolved authority, unavailable data, an unwilling user group, a weak outcome, or an operating commitment nobody wants to own.
These outcomes preserve speed. They prevent a team from spending its short delivery window discovering that the real blocker was visible before coding began.
A small initiative earns attention when it can close a loop: from a meaningful problem, through a bounded intervention, into real use, measurable evidence, and a decision. The code may be finished quickly. The value exists only when the surrounding work is ready to absorb it.