Claude Partner Network memberAI, engineered for the enterprise
Strategy

How to scope an AI proof of concept you can actually judge.

The demo goes well. Everyone in the room nods. Three months later the pilot is still "in progress," the vendor is asking for an extension, and nobody can say whether the thing worked. That outcome was decided on day one, when nobody wrote down what success would look like.

PoCs rarely fail in the model. They fail in the scoping.

A proof of concept exists to answer one question: should we invest more? Most of them can't answer it. They start with a broad ambition ("see what AI can do for claims"), run on examples the vendor picked, and end with a demo that looks good and proves nothing. The decision to continue gets made on impressions, because impressions are all anyone collected.

The fix is not a better model. It is a scope tight enough that on a known date, a named person can look at agreed evidence and say yes, no, or not yet. Everything below serves that one meeting.

Five things a judgeable PoC has

One workflow means one. Not "customer service," but "first-response drafts for coverage questions on policy type X." A narrow workflow lets you build a real evaluation set and gives the result a clear owner. If the PoC works, you scale sideways to adjacent workflows. If it fails, you know exactly what failed.

The evaluation set is the piece most teams skip, and it is the piece that makes judgment possible. Pull real cases from the last quarter, including the ugly ones: the ambiguous intake, the document with a missing page, the request that should have been escalated. Freeze the set before the build starts. If the vendor can add or remove cases mid-flight, you are grading their homework with their answer key.

Kill criteria matter as much as success criteria. Decide in advance what result means stop: an accuracy floor on the frozen set, a category of error you cannot tolerate at any frequency, a review burden that exceeds the work it replaces. Writing these down turns a kill from an awkward political event into a scheduled checkpoint.

1One workflowa single named process with a single owner, not a department
2One outcome metricthe number that decides, agreed before any build begins
3Fixed eval setreal historical cases, frozen on day one, edge cases included
4Decision datea calendar date when someone says yes, no, or not yet
5Kill criteriathe results that mean stop, written down in advance
The five elements that make a proof of concept judgeable. Remove any one and the decision meeting becomes a debate.

What to give the vendor on day one

A vendor can only be held to a standard you supply the materials for. If you show up with an idea and no artifacts, expect a demo built on synthetic examples and a result you cannot verify. The day-one package is short but non-negotiable.

  • Real cases. A frozen set of historical examples from the actual workflow, redacted where required, with the correct outcome recorded for each. Include the failures and edge cases, not just the clean runs.
  • The current baseline. How the work is done today, who does it, and what a mistake costs. Without a baseline, "the AI got it right" has nothing to be compared against.
  • An access-control map. Which systems the PoC may touch, which data it may see, and who approves exceptions. Deciding this after the build is how shadow integrations happen.
  • A named reviewer. One person with the expertise to judge outputs and the authority to say no. A committee reviewing occasionally is not oversight, it is diffusion of responsibility.
  • The decision date. On the calendar, with the decision-maker's name next to it, before any work begins.

What the vendor owes you back

The exchange runs both ways. A vendor who accepts your frozen eval set and your decision date should hand back evidence, not a highlight reel. Ask for these deliverables in writing before the engagement starts, and treat reluctance as data.

  • The written criterion, first. The success metric and kill criteria, signed off before any build starts. If the vendor wants to "figure out the metric as we go," the PoC is already unjudgeable.
  • Results on your set, not theirs. Performance on the frozen cases, with the misses listed alongside the hits. A miss list with explanations tells you more about production readiness than any accuracy figure.
  • A log of every output. Each response tied to its inputs, its sources, and whether a human accepted or corrected it. If the PoC has no audit trail, the production system won't either.
  • A data-flow note. Where your data went, where it is stored, and confirmation it was not used to train anyone's models. One page is enough. Silence is not.
  • An honest path to production. What breaks at real volume, what governance work remains, and roughly what the next phase requires. A vendor who says "just flip it on" has not thought about it.

The decision meeting has three outcomes, and one of them is kill

On the decision date there are exactly three moves. Scale: the criterion was met on the frozen set, the reviewer trusts the outputs, and the next phase gets funded. Iterate once: the results are close, the failure pattern is specific and fixable, so you set one new date and one new build cycle, not a rolling extension. Kill: the criteria were not met, and you stop.

A kill on schedule is the PoC working as designed. You spent a bounded amount to learn that this workflow, this data, or this approach is not ready, and you learned it before committing a budget line and a headcount. The pilots that damage AI programs are not the ones that end. They are the ones that never do.

This is also how you should evaluate anyone offering to run a PoC for you, including us. Our Governed Pilot exists because we think the criterion belongs in the agreement itself: fixed scope, a success measure agreed up front, and if it isn't met, you don't pay. Whoever you work with, insist on that structure. A vendor confident in their system will accept it. A vendor who resists is telling you how the pilot will end.

A PoC killed on schedule is a cheap answer. A pilot that never ends is an expensive way to avoid one.

You do not need a bigger model or a longer timeline to get a proof of concept you can judge. You need one workflow, one metric, a frozen set of real cases, a date on the calendar, and the nerve to write down what stop looks like. Everything else is a demo.

Take the readiness assessment How we help