Every implementation partner has a polished demo and a slide about responsible AI. Neither tells you whether they can put Claude in front of your regulators, your auditors, or your general counsel and survive the questions that follow. These eight questions will. They apply to us too, and we say so below.
A demo proves the model works. It proves nothing about whether the partner can operate it inside your control environment. So skip the feature tour and open with the two questions that expose whether governance is a document or a habit.
The first is simple: show me the framework. A production-grade partner has a written governance framework covering data protection, human oversight, auditability, and access control, and they can hand it to you before contract signature. Ours is called Governed-by-Design, and any prospect can read it. If a partner's answer is a slide with the word 'responsible' on it, you have learned what you need to know.
The second question is about people. Sales engineers build demos. Delivery engineers build systems. Ask who will actually write the code, run the evaluations, and sit in your architecture reviews, and ask to meet them before you sign. If the people in the pitch will not be the people in the engagement, discount everything the pitch promised.
Open-ended time-and-materials engagements reward the partner for the project never finishing. That is not cynicism, it is incentive design. The third question forces the issue: is the first engagement fixed in scope, with a written success criterion, and what happens commercially if the criterion is not met? Our Governed Pilot answers this with a guarantee, and if the agreed criterion is not met, the client does not pay. You should not accept less from anyone, including us. A partner unwilling to define success in writing is telling you they do not expect to hit it.
The fourth question is where most vendors fall apart: how do you evaluate the system? Not vibes, not a handful of cherry-picked prompts. Ask what the evaluation set looks like, who writes the test cases, how many there are, how regressions are caught when a prompt or model version changes, and whether the eval harness is delivered to you as an artifact. A partner who has actually shipped to production will answer in specifics, because they have been burned by the alternative.
A partner who cannot describe their evaluation discipline in concrete detail has never had a system challenged by an auditor.
In a regulated environment, an answer without a source is a liability. Question five: how does the system ground its outputs, and can every claim be traced to a specific document, page, or record? Retrieval-augmented architectures make this possible, but only if citations are built in from the first sprint rather than bolted on before go-live. Ask to see a cited answer from a real system, then ask what happens when the retrieval layer finds nothing. Silence and refusal are the right behaviors. Confident invention is the failure mode you are paying to avoid.
Question six is the one your compliance team will eventually ask, so ask it now: what does the audit trail capture? A serious answer covers the prompt, the retrieved sources, the model version, the output, the timestamp, and the identity of the human who reviewed or approved the result. Ask where those logs live, who can read them, how long they are retained, and whether the partner can reconstruct any individual answer six months later. If the audit story is 'the platform logs things', keep asking until you get to fields and retention periods.
Question seven is about what you own when the engagement ends. Some partners deliver a system. Others deliver a dependency. Ask whether you receive the prompts, the retrieval configuration, the evaluation sets, the infrastructure code, and the documentation needed for your own team to run and modify everything without them. Ask whether anything sits on the partner's own hosted platform, and what migration off it looks like. The right answer is that the entire build lives in your environment from day one and their exit costs you a handover meeting, not a rebuild.
The eighth question is the honesty test, for the partner and for us. Ask for references from regulated deployments and listen carefully to the answer. We do not publish named clients, and any partner working in healthcare, legal, or financial services will face the same constraint, because those clients rarely permit public naming. But confidentiality is not a reason to offer nothing. A credible partner can arrange direct reference conversations under NDA, or walk you through anonymized architectures, evaluation results, and governance artifacts in enough depth that the work is obviously real. Vague gestures at 'major institutions' with nothing behind them deserve the skepticism you are already feeling.
Run all eight questions in the same meeting and patterns emerge fast. Deck-sellers answer in adjectives. Builders answer in artifacts: a framework document, an eval harness, a log schema, a handover plan. You are not scoring charisma. You are checking whether the partner has already solved the problems you are about to have.
One more suggestion. Send the questions in advance. A partner who welcomes that is confident in their answers. A partner who wants to keep the conversation on the demo is telling you where their strength ends.
We built AI Definitive to pass this list, from the Governed Pilot's success guarantee to the Governed-by-Design framework we hand over before signature. If you want to test that claim, the free Claude Readiness Assessment is a reasonable place to start, and these eight questions are a reasonable thing to bring to it.