A claims assistant drafts a denial. An adjuster accepts it and the letter goes out. From that moment, no regulator cares that a model wrote the first draft. The carrier owns the decision, the reasoning, and the paper trail. That single fact should shape every AI deployment in claims and underwriting.
In most industries, AI assists a workflow that ends somewhere else. In insurance, the workflow ends in a regulated act: a coverage determination, an eligibility call, a price. When a model drafts that act, the distance between the model's output and the regulated decision can be one click. That is why the citation trail, the review gate, and the decision log are not compliance decoration layered on top of a claims assistant. They are the difference between a tool you can defend and one you have to unplug mid-exam.
State regulators have started saying this out loud. Many insurance departments have adopted versions of the NAIC's model bulletin on insurers' use of AI, which expects carriers to maintain a written program governing AI systems, including internal governance, testing, and oversight of third-party models and data. Adoption and wording vary by state, and some states have gone further with their own statutes or rules. Map the expectations across your licensed states rather than building to a single assumed standard.
A claims assist that says "water damage, not covered" is useless and dangerous in equal measure if it cannot point to the exclusion it relied on. The bar is specific: the form number, the edition date, the endorsement that modified it, the paragraph. That means retrieval has to run against the actual policy forms in force for that policy on that date of loss, not a generic library of similar forms. Endorsements change everything, and a citation to the wrong edition is worse than no citation, because it looks like diligence.
Denials need a second gate. The model drafts, and an adjuster with the right authority level decides. The log records both: which passages the model cited, and which human reviewed and approved. Underwriting referrals work the same way. A model can flag, score, and draft. A licensed underwriter declines. That division of labor is what you will point to when someone asks who made the decision.
When a decision relies on consumer report data, federal adverse action duties under the FCRA apply regardless of what generated the recommendation. Beyond that, most states' unfair claims settlement practices laws require carriers to explain the basis of a denial with reference to the policy provisions relied on. And a growing number of states have enacted or proposed rules specifically addressing insurers' use of algorithms and external data in particular lines, some with quantitative testing expectations for unfair discrimination.
The design consequence is straightforward. Every automated recommendation that could become an adverse decision needs reasons a human can read, defend, and put in a letter. Reasons tied to filed rating factors and actual policy provisions, not a confidence score. If your footprint spans states with different expectations, build to the strictest one. Running two explainability standards is more expensive than running one good one.
Market conduct exams run on evidence, and the examiner's questions about AI are predictable versions of questions they already ask. What systems touched this decision? Who approved this denial? Show me your testing. Show me how a complaint traces back to the file. If your AI deployment cannot answer those from logs, you will be reconstructing answers by hand under deadline, which is where findings come from.
The evidence pack is buildable in advance, and building it in advance is dramatically cheaper than assembling it during an exam.
Foundation models get updated. Prompts get tuned. Policy forms get refiled and your retrieval corpus shifts with them. Any of these can move outcomes on the same claim file, and a carrier that cannot say which version of the system made a given decision cannot defend that decision later. Treat a model change the way you treat a rate change internally: documented, tested against a known baseline, and approved before it touches production files.
The baseline is a frozen evaluation set of adjudicated claims and underwriting files with known-correct outcomes. Re-run it on every change. And watch your adjusters: a rising override rate, where humans increasingly reverse the model's drafts, is usually your earliest drift signal, well before any dashboard metric moves.
This is the shape of AI Definitive's Governed-by-Design work in regulated lines: citations, human review, logs, and change control built in from the first week, not retrofitted before an exam. If you want to know where your current claims or underwriting workflow stands, the free Claude Readiness Assessment is the fastest honest answer.