We use cookies to understand how this site is used. Privacy policy

    Skip to main content
    Kriv AI

    Insurance AI Governance

    Why Is NAIC AI Model Governance So Hard to Implement for AI?

    Compliance and AI teams that build a good-faith governance program to the letter of the NAIC bulletin still walk into exam findings. Here is what actually makes implementation hard, with the regulatory record behind it.

    NAIC AI model governance is hard to implement because more than twenty states have each adopted the bulletin with their own variations, carriers must independently document every third-party AI vendor rather than rely on the vendor, and the evidentiary bar keeps shifting as NAIC's evaluation tool moves through an active 12-state pilot.

    cause

    The Real Reasons NAIC AI Model Governance Is Hard to Implement

    One Bulletin, More Than Twenty Different Compliance Programs

    The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers was adopted by NAIC membership in December 2023 as a single template. In practice it never stayed one template. By late 2025, 23 states and Washington, D.C. had adopted the bulletin in some form, and adoption has kept accelerating through 2026. The problem for a compliance team is that not every state adopted it verbatim. States including Colorado, Virginia, Connecticut, and Pennsylvania have moved toward more prescriptive requirements layered on top of the bulletin's baseline, and Colorado's own regulation under SB 21-169 already imposes separate testing and attestation obligations on life, auto, and health benefit insurers that go beyond what the bulletin itself requires.

    For a carrier licensed in a dozen or more states, that means a single AI governance program has to satisfy a dozen or more slightly different sets of expectations at once, not one shared standard. A governance document written to the bulletin's general language can still leave a carrier short in a state that added its own documentation, testing, or attestation layer, and there is no single master checklist that covers every variant. That is the first reason implementation is hard: the target is not fixed, and it is not the same target in every state where the carrier operates.

    Third-Party AI Vendors Do Not Reduce Your Accountability

    Most insurers do not build their underwriting, claims, or fraud-detection AI in-house. They license it. That creates the second implementation problem: outsourcing the AI does not outsource the compliance obligation. The 2023 Model Bulletin already required carriers to apply the same governance rigor to third-party models as to internally built ones, and the NAIC's advancing proposal for a third-party AI vendor registry does not change that. Registration, where it exists, creates regulatory visibility into a vendor's model, not a safe harbor for the carrier that uses it. A carrier that uses a registered model is still answerable for the model's behavior.

    In practice, this means a compliance team has to go back to every AI vendor and ask for training data sources, validation records, bias testing methodology, and known limitations, and many vendors simply cannot produce that documentation on request. When a vendor cannot provide model documentation, validation records, or audit support, that gap becomes the carrier's problem during an exam, not the vendor's. Reconstructing that vendor evidence retroactively, once an examiner has already asked for it, takes far longer and costs far more than building it into the vendor contract and onboarding process from the start.

    The Evidentiary Bar Is Still Moving Through an Active Pilot

    The third reason implementation is hard is that the standard itself is not finished. The NAIC's AI Systems Evaluation Tool, the structured framework examiners will use to assess an insurer's AI governance during market conduct and financial exams, is currently in a multistate pilot running from March through September 2026 across 12 states: California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia, and Wisconsin. The NAIC plans to revise the Tool based on pilot feedback in September and October 2026, re-expose it for public comment, and consider it for adoption at the NAIC's Fall National Meeting in November 2026.

    That means any governance program built today is being built against a standard that is still being field-tested by regulators themselves. A carrier that documents exactly what today's pilot version of the Tool asks for could still see the bar move once the revised version is adopted later in 2026. Building a program that is directionally right, covering governance structure, model inventory, validation, bias testing, third-party oversight, and consumer-impact assessment, matters more than optimizing for the exact wording of a tool that is still being revised.

    Bias Testing Requires Infrastructure Most Carriers Have Not Built

    The fourth reason is technical, not just procedural. The bulletin and the state laws layered on top of it expect quantitative bias testing, not a policy statement that the carrier cares about fairness. That means measuring disparate impact in model outputs across protected classes, checking calibration by subgroup, and testing for proxy variables such as zip code or credit attributes that can reproduce a prohibited bias even when the model never uses a protected characteristic directly. Building and running that testing pipeline, and keeping a written record of the data used, the metric values computed, and who reviewed the result, is a data science and compliance function most carriers have never had to build together before.

    The gap is measurable. Despite 92% of health insurers reporting current or planned AI use, nearly one-third of health insurers still do not regularly test their models for bias or discrimination. That gap between AI adoption and AI testing is exactly what an examiner is now trained to look for, and it is why so many carriers that believe they are compliant are not, once someone actually asks for the test results.

    evidence

    What the Regulatory Record Shows

    The Model Bulletin itself is not a draft proposal. NAIC membership approved it in December 2023 after two public comment periods, and the NAIC has been explicit that the goal is to set clear expectations for AI governance while balancing innovation against the risks specific to AI-driven decisions. Adoption has moved from a handful of early states, Alaska was first in February 2024, to a majority of the country inside two and a half years, and state insurance departments in non-adopting states are increasingly applying the same expectations informally through routine market conduct examinations even without a formal bulletin adoption.

    The NAIC's AI Systems Evaluation Tool pilot is the clearest evidence that regulators are actively operationalizing the bulletin, not just publishing it. The pilot's four exhibits ask a carrier to quantify its AI usage, describe its governance and risk-assessment structure, detail any high-risk AI systems, and document the underlying data, which mirrors almost exactly the documentation gaps carriers report struggling with: inventory, ownership, testing evidence, and vendor oversight. A revised Tool is expected to move toward formal adoption at the NAIC's Fall 2026 National Meeting, which means the current pilot period is the last checkpoint before this becomes a standard, not experimental, part of exam scope.

    On the vendor side, the NAIC's proposed third-party AI model and data registry is moving on a parallel track, with framework exposure expected in the third quarter of 2026 and possible adoption consideration in November 2026. Its own stated purpose is to give regulators visibility into vendor models, explicitly without shifting the carrier's underlying accountability for how those models are used.

    framework

    How to Build NAIC-Aligned AI Governance That Survives an Exam

    Start with a real inventory, not a partial one. List every AI or algorithmic model that touches underwriting, rating, claims, fraud detection, or marketing, whether it was built internally or licensed from a vendor, and name an accountable owner for each. If a model cannot be tied to a named owner and a documented validation record, that is the first gap an examiner will find, and it is the cheapest one to fix before an exam rather than during one.

    Next, build the vendor file before you need it, not after an examiner asks. For every third-party AI tool, request in writing the vendor's training data sources, testing methodology, known limitations, and change-management process, and keep that file current as the vendor updates its model. Treat a vendor's refusal or inability to provide this documentation as a governance finding of its own, not a reason to skip the question.

    Then run the quantitative bias tests regulators actually look for: disparate impact across protected classes on real decision outcomes such as denial rate, pricing, or claim payout, calibration checks by subgroup, and a review of proxy variables like zip code that can reintroduce bias indirectly. Document the metric, the threshold, the result, and who reviewed it, in a form that would make sense to someone who was not in the room when the test ran. Finally, build the program to the direction of the bulletin and the pilot Evaluation Tool rather than chasing the exact wording of either, since both are still moving through 2026, and a program built on the underlying principles will hold up better than one built to match a specific document that is about to be revised.

    differentiation

    How Kriv AI Helps

    Kriv AI builds the specific artifacts NAIC's evaluation framework and state examiners ask for: a named-owner AI model inventory, vendor due diligence files for third-party AI, documented bias testing with the metric and reviewer trail examiners expect, and a governance program mapped to the bulletin's baseline plus the state-specific layers that apply to where a carrier is licensed. Work for regulated insurers is billed at Kriv AI's standard regulated-industry rate of $200 per hour, with fractional AI governance lead engagements at $300 to $400 per hour for carriers that need ongoing oversight through the rest of the NAIC pilot period rather than a one-time review. All engagements carry an $8,000 minimum.

    If a governance program is already built but has never been tested against what an examiner would actually ask for, a discovery call is the fastest way to find the gaps before a regulator does.

    Straight answers

    Frequently asked questions about Why Is NAIC AI Model Governance So Hard to Implement for AI?

    Why is NAIC AI model governance so hard to implement?

    It is hard because more than twenty states have adopted the NAIC Model Bulletin with their own variations, carriers remain fully accountable for third-party AI vendors even when the vendor cannot produce documentation, and the NAIC's own AI Systems Evaluation Tool is still moving through an active 12-state pilot, so the exact evidentiary bar keeps shifting.

    Does every state require the same NAIC AI governance program?

    No. States including Colorado, Virginia, Connecticut, and Pennsylvania have added more prescriptive requirements on top of the bulletin's baseline, and Colorado's SB 21-169 imposes separate testing and attestation duties for life, auto, and health benefit insurers. A carrier licensed in multiple states needs a program that satisfies the strictest applicable requirement, not just the bulletin's general language.

    Are we responsible for AI models we license from a vendor?

    Yes. The NAIC Model Bulletin requires carriers to apply the same governance rigor to third-party AI as to models built in-house, and the NAIC's proposed vendor registry explicitly does not shift accountability away from the carrier. A carrier using a registered or well-known vendor model is still answerable for how that model behaves.

    What is the NAIC AI Systems Evaluation Tool pilot?

    It is the structured framework state examiners are testing to assess insurer AI governance during market conduct and financial exams. The pilot runs March through September 2026 across 12 states, and the NAIC plans to revise the Tool based on pilot results and consider it for formal adoption at the Fall 2026 National Meeting.

    What documentation should we have ready before an examiner asks?

    A named-owner model inventory, pre-deployment validation records, documented bias testing results with the metric and reviewer noted, a change log for model updates, and a vendor due diligence file for every third-party AI tool in use, covering training data sources, testing methodology, and known limitations.

    How much does it cost to build NAIC-aligned AI governance with Kriv AI?

    Kriv AI's regulated-industry work starts at $200 per hour for standard engagements and $300 to $400 per hour for fractional AI governance lead roles, with an $8,000 minimum engagement. Scope and cost depend on how many models and vendors are in scope and how much governance documentation already exists.

    Talk to the team that would do the work

    Bring your requirements to a working session with the person who'll actually deliver.

    Book a Discovery Call