Select Services Partner, Claude Partner Network•AI, engineered for the enterprise
◆Insurance

The written program is what the regulator reads first so build it like it will be read.

When a market conduct exam or a regulator inquiry arrives, one of the earliest requests is now predictable: produce your written AI governance program. Not the model, not the vendor deck. The document. State insurance regulators have converged on this expectation, and that convergence changes what compliance work looks like.

The document itself is now the compliance surface

For years, insurers could treat AI oversight as a set of good habits distributed across actuarial, IT, and compliance teams. That era is closing. Regulators in a growing number of states expect a single written artifact that describes how the company governs AI systems affecting consumers, and they expect to be able to request it. An examiner who receives a two-page memo draws a conclusion before reading a single model file.

This flips the usual sequencing. The model is not audited first. The program is. A well-built program document does two jobs at once: it satisfies the request, and it gives the examiner a map that keeps the review inside boundaries you drew. A missing or thin program invites the examiner to draw the map themselves.

What belongs in the program

The specific wording varies by state, but the components regulators look for have stabilized. A program that covers the following, with real content behind each heading rather than aspirational language, will survive most first-pass reviews.

  • Scope and inventory. A definition of which systems the program covers, and a maintained register of every AI system in use, including models purchased inside vendor platforms.
  • A named accountable owner. A specific person, by role, who owns the program and answers for it. Committees can advise; a committee cannot be the accountable party.
  • Risk tiering. A method for sorting systems by consumer impact, so a claims-denial model gets more scrutiny than a mailroom classifier, and the reasoning is written down.
  • Pre-deployment and ongoing testing. What gets tested before a system touches a consumer decision, what gets monitored after, on what cadence, and who reviews the results.
  • Documentation and record retention. Where test results, decision logs, and model changes live, how long they are kept, and how they would be produced for an exam.
  • Vendor oversight. How third-party models are evaluated before purchase and monitored after, since the regulator holds the insurer accountable regardless of who built the model.
  • Consumer recourse. A defined path for a policyholder to contest a decision an AI system influenced, including who performs the human review and how the outcome is recorded.

Building the inventory when nobody knows what is deployed

The inventory is where most programs stall, because at most carriers no single person knows what is running. Underwriting bought a scoring service years ago. A claims platform added a fraud model in a routine vendor upgrade. A regional team adopted a triage tool that never crossed a procurement desk. All of it is in scope.

The reliable path starts with money, not with asking people what AI they use. Pull vendor contracts and procurement records and flag anything that scores, ranks, predicts, or classifies. Then interview the teams closest to consumer decisions, because the tools they rely on daily often never appear in a contract line item. Each entry in the resulting register needs a small set of fields: what the system does, what data feeds it, which consumer decisions it touches, who owns it, and its risk tier. A register without a named owner per system is a list, not an inventory.

Treat the register as a living record with a review cadence, because the version that matters is the one that exists on the day the exam letter arrives.

Contract pullProcurement scanTeam interviewsDraft registerOwner sign-offRisk tiering
Building the system inventory from records first, self-reporting second.

The testing question that carries the most weight

One question separates programs that hold up from programs that read well: can you show the outcome distribution of your AI-influenced decisions across protected classes, and can you explain what you see? Everything else in the program supports this. If the answer is no, the rest is scaffolding around an empty room.

This is genuinely hard for insurers, and honest programs say so. Most carriers do not collect protected class data directly, which means measuring outcomes often requires inference methods, and choosing an inference method is itself a governed decision that belongs in your documentation. The mechanics matter less than the posture: a defined methodology, results recorded before deployment and on a schedule after, and a written review when the distribution shifts. An examiner is far more persuaded by a documented method with acknowledged limitations than by a claim of fairness with nothing behind it.

Explanation is the second half of the requirement. A distribution alone proves you looked. The program also needs to show who reviewed it, what threshold triggers escalation, and what happened the last time something looked wrong. Testing without a documented response path is monitoring theater.

If you cannot produce the outcome distribution across protected classes for a model that touches claims or underwriting, you do not have a testing program yet. You have a plan for one.

Where to start

Sequence matters. Draft the program skeleton and name the accountable owner first, because both are prerequisites for everything else. Build the inventory second. Tier the systems third, and point your first real testing effort at the highest-tier system rather than trying to test everything at once. A program that governs one consequential model well is more defensible than one that governs twenty models on paper.

This is where we work. AI Definitive is a Claude specialist and a Select Services Partner in the Claude Partner Network, serving regulated industries, and our Governed-by-Design framework maps directly onto what insurance regulators ask for: human oversight, auditability with cited and logged outputs, access controls, and data protection with client data never used to train models. Our Model Governance & Audit Pack is built to produce the documentation layer described above, and our engagements are designed to support the regulatory frameworks insurers answer to. A free Claude Readiness Assessment is a reasonable first step; a Governed Pilot with a fixed scope and a success guarantee is the second, and if the agreed criterion is not met, you do not pay.

The regulators asking for a written AI governance program are not asking for a promise. They are asking for an artifact that proves the promise was operationalized. Build the document so that every section points to something real, and the exam becomes a production exercise instead of a scramble.

Take the readiness assessment How we help