AIStrategyProduction

The POC-to-Production Checklist: 12 Questions Before You Greenlight an AI Pilot

A proof of concept tells you an idea is possible. It says nothing about whether it's deployable. These are the twelve questions that surface the gap, before you spend six months and a budget discovering it the hard way.

Every AI project that ever stalled started life as a demo that worked. The model answered the question, the room nodded, and the budget got approved. The trouble is that a demo and a production system are two different animals, and the distance between them is exactly where most enterprise AI quietly goes to die.

We wrote about that distance in The AI POC Graveyard, the roughly 80% of pilots that never reach a single real user. This is the practical companion: a pre-flight checklist you can run before you greenlight a pilot, so the deployability gaps surface in a planning meeting instead of in month six.

How to use this

Ask these twelve questions of any AI pilot before it gets budget. If you can't answer one of them with something better than "we'll figure it out later," you've just found the risk that will stall the project. A confident "we don't know yet" is a finding, not a failure, it tells you exactly where to look first.

Data & integration: the foundation nothing works without

More pilots die here than anywhere else, and almost none of it shows up in a demo built on clean, curated data.

  1. Where does the data actually live, and who owns it? Production data is scattered across CRMs, ERPs, and siloed spreadsheets, each with its own owner and quirks. If you can't name the systems and the people, you can't estimate the work.
  2. Is the demo running on hand-picked data, or the real thing? A prototype on curated examples proves the idea. It says nothing about how the system behaves on the messy, inconsistent, half-empty records that make up real operations.
  3. What must this integrate with, and what are its limits? API rate limits, authentication, and data-security rules on legacy systems routinely cost more engineering than the model itself. Map them now, not after the model is built.

Reliability & evaluation: can you actually trust the output?

A demo is a scripted best case. Production is every case, including the ones no one anticipated.

  1. Do you have an eval set, or just a good feeling? You can't ship what you can't measure. Without a labelled set of examples and a defensible quality bar, "it works" is a vibe, not an engineering claim.
  2. What happens on the edge case the demo never saw? Real users paste malformed input and ask the unplanned question. Define fallbacks, hard limits, and safe failure modes up front, they are the product, not an afterthought.
  3. How will you catch wrong answers before a customer does? Probabilistic systems hallucinate and invent policies that don't exist. Decide how you'll detect that, through guardrails, human review, or monitoring, before launch rather than after the first incident.
Notice how few of these questions are about the model. That's the point.

Operations & ownership: who runs it after launch?

Shipping the model is the start of the work, not the end. Someone has to keep it alive.

  1. Who owns this in production at 2am when it breaks? If the answer is "the team that built the demo, in their spare time," the pilot has no real operational owner, and unowned systems rot.
  2. How will you monitor drift, latency, cost, and quality? Models degrade as inputs shift and the world changes. Decide what goes on the dashboard, and who watches it, before you depend on the system.
  3. What's the rollback plan? Every production system needs a kill switch and a way back to the last known-good state. If there's no way to turn it off safely, it isn't ready to turn on.

Economics & scope: is it worth building at all?

AI has unit economics. Sometimes the honest answer is that the LLM is the wrong tool for the job.

  1. What's the fully-loaded cost per call? Tokens are the visible bill. Infrastructure, evals, maintenance, and human oversight are the rest. Weigh value-per-call times volume against that total, not against the API price alone.
  2. Could a rule, script, or lookup do this better? For plenty of jobs, deterministic code beats an LLM on cost, speed, and reliability. Reserve the model for the parts that genuinely need it.
  3. What's the smallest slice you can ship, and what does "good enough" mean in numbers? Define the thin first release and the specific metric that clears it for launch. "Greenlight" should be a number you agreed on in advance, not a feeling in a review meeting.

12

questions that decide whether a pilot ships, or becomes another headstone in the POC graveyard.

The checklist, in one screen

The twelve questions, stripped to their essentials. Screenshot this and take it into your next pilot review.

  1. Where does the data live, and who owns it?
  2. Is the demo on hand-picked data, or the real thing?
  3. What must it integrate with, and what are the limits?
  4. Do you have an eval set and a quality bar?
  5. What happens on the unplanned edge case?
  6. How will you catch wrong answers before a customer does?
  7. Who owns it in production at 2am?
  8. How will you monitor drift, latency, cost, and quality?
  9. What's the rollback plan?
  10. What's the fully-loaded cost per call?
  11. Could a rule or script do it better?
  12. What's the smallest slice to ship, and what clears it?

Greenlighting an AI pilot well is a product and operations decision, not just a technical one. It's about knowing what you're building, who runs it, and what "done" means, before the budget is spent.

Answer these twelve honestly and you won't eliminate every risk. But you will have moved the expensive surprises out of month six and into the room where they're still cheap to fix.

Tags AI
Share
Reef TRH
AI Architecture & Production Engineering

We turn fragile AI proofs of concept into stable, production ready systems, bridging engineering and operations so your AI actually ships and survives production.

Contact us

Not sure your pilot is ready to greenlight?

Book a strategy call No pitch deck. A 30 minute working conversation about your specific bottleneck.