The demo goes well. Everyone in the room nods. Three months later the pilot is still "in progress," the vendor is asking for an extension, and nobody can say whether the thing worked. That outcome was decided on day one, when nobody wrote down what success would look like.
A proof of concept exists to answer one question: should we invest more? Most of them can't answer it. They start with a broad ambition ("see what AI can do for claims"), run on examples the vendor picked, and end with a demo that looks good and proves nothing. The decision to continue gets made on impressions, because impressions are all anyone collected.
The fix is not a better model. It is a scope tight enough that on a known date, a named person can look at agreed evidence and say yes, no, or not yet. Everything below serves that one meeting.
One workflow means one. Not "customer service," but "first-response drafts for coverage questions on policy type X." A narrow workflow lets you build a real evaluation set and gives the result a clear owner. If the PoC works, you scale sideways to adjacent workflows. If it fails, you know exactly what failed.
The evaluation set is the piece most teams skip, and it is the piece that makes judgment possible. Pull real cases from the last quarter, including the ugly ones: the ambiguous intake, the document with a missing page, the request that should have been escalated. Freeze the set before the build starts. If the vendor can add or remove cases mid-flight, you are grading their homework with their answer key.
Kill criteria matter as much as success criteria. Decide in advance what result means stop: an accuracy floor on the frozen set, a category of error you cannot tolerate at any frequency, a review burden that exceeds the work it replaces. Writing these down turns a kill from an awkward political event into a scheduled checkpoint.
A vendor can only be held to a standard you supply the materials for. If you show up with an idea and no artifacts, expect a demo built on synthetic examples and a result you cannot verify. The day-one package is short but non-negotiable.
The exchange runs both ways. A vendor who accepts your frozen eval set and your decision date should hand back evidence, not a highlight reel. Ask for these deliverables in writing before the engagement starts, and treat reluctance as data.
On the decision date there are exactly three moves. Scale: the criterion was met on the frozen set, the reviewer trusts the outputs, and the next phase gets funded. Iterate once: the results are close, the failure pattern is specific and fixable, so you set one new date and one new build cycle, not a rolling extension. Kill: the criteria were not met, and you stop.
A kill on schedule is the PoC working as designed. You spent a bounded amount to learn that this workflow, this data, or this approach is not ready, and you learned it before committing a budget line and a headcount. The pilots that damage AI programs are not the ones that end. They are the ones that never do.
This is also how you should evaluate anyone offering to run a PoC for you, including us. Our Governed Pilot exists because we think the criterion belongs in the agreement itself: fixed scope, a success measure agreed up front, and if it isn't met, you don't pay. Whoever you work with, insist on that structure. A vendor confident in their system will accept it. A vendor who resists is telling you how the pilot will end.
A PoC killed on schedule is a cheap answer. A pilot that never ends is an expensive way to avoid one.
You do not need a bigger model or a longer timeline to get a proof of concept you can judge. You need one workflow, one metric, a frozen set of real cases, a date on the calendar, and the nerve to write down what stop looks like. Everything else is a demo.