We use cookies to understand how this site is used. Privacy policy

    Skip to main content
    Kriv AI

    Financial Services Governance

    Model Risk Management for AI: From SR 11-7 to SR 26-2

    How the shift from SR 11-7 to SR 26-2 changes model validation, inventory, and governance for generative and agentic AI.

    context

    What SR 26-2 Changes

    Why Generative AI and LLMs Broke SR 11-7

    SR 11-7, the Federal Reserve's 2011 model risk management guidance, was written for statistical and machine learning models with stable inputs and measurable accuracy. A large language model or an agentic AI system doesn't backtest the same way: its behavior can vary by prompt, its outputs are generative rather than a single predicted value, and a single deployment can span dozens of use cases rather than one model, one purpose.

    Side-by-Side: What SR 11-7 Assumed vs. What SR 26-2 Addresses

    SR 11-7 assumed a model has a fixed purpose, a measurable output, and a backtestable accuracy metric. SR 26-2 extends that framework to cover generative and agentic systems where validation has to focus on output quality sampling, guardrail testing, and use-case-specific risk tiering rather than a single accuracy number.

    framework

    Three Lines of Defense, Reinterpreted for AI

    The first line (the business unit deploying the model), second line (independent risk and compliance oversight), and third line (internal audit) structure still applies, but AI adds new questions to each line: who owns a prompt-engineering change to a production LLM, who validates an agent's tool-use permissions, and who audits a vendor-supplied foundation model your team doesn't control the training data for.

    validation

    Model Validation for LLMs and Agentic AI

    Why Backtesting Fails on Generative Models

    A generative model doesn't produce one measurable prediction to check against ground truth. Validation instead needs structured output sampling against a rubric, red-teaming for guardrail failures, and ongoing drift monitoring on output quality rather than a single input-output accuracy check.

    inventory

    Model Inventory and Risk Tiering

    Every AI system in production, including third-party vendor tools with an embedded LLM, needs to be in a model inventory with a risk tier assigned based on the decision it influences: a customer-facing chatbot answering general questions sits in a different tier than an agent that can initiate a wire transfer or approve a loan.

    governance

    Governance, Roles, and Board Reporting

    SR 26-2 expects board and senior-management reporting on AI model risk with the same rigor as traditional model risk, meaning a bank or asset manager needs a defined owner for AI governance, not an informal arrangement where whoever deployed the model also validates it.

    roadmap

    A 90-Day Compliance Roadmap

    A typical path: weeks one through three, inventory every AI system currently in production or pilot, including vendor tools with embedded AI. Weeks four through six, risk-tier each one and identify validation gaps against SR 26-2 expectations. Weeks seven through twelve, close the highest-risk gaps first and stand up the ongoing governance structure to keep the inventory current.

    engagement

    How Kriv AI Implements This

    A regulated financial services client engaged Kriv AI to build an AI model inventory and validation framework ahead of an SR 26-2 gap assessment, the kind of anonymized engagement pattern we bring to a scoped model risk management project: inventory, risk tiering, validation framework design, and an ongoing governance structure your board can actually report against.

    Straight answers

    Frequently asked questions about Model Risk Management for AI: From SR 11-7 to SR 26-2

    What is SR 26-2?

    SR 26-2 is the Federal Reserve's updated model risk management guidance, extending the SR 11-7 framework to explicitly address generative AI, large language models, and agentic AI systems that don't validate the same way traditional statistical models do.

    Does SR 11-7 still apply, or has it been replaced?

    SR 26-2 builds on SR 11-7's three-lines-of-defense structure rather than discarding it, extending the validation and governance expectations to cover generative and agentic AI specifically.

    Does SR 11-7 apply to large language models?

    The original SR 11-7 guidance predates LLMs and doesn't address them explicitly, which is exactly the gap SR 26-2 is designed to close with generative-AI-specific validation expectations.

    How do you validate an LLM if you can't backtest it like a traditional model?

    Through structured output sampling against a rubric, red-teaming for guardrail failures, and ongoing drift monitoring on output quality, rather than a single input-output accuracy metric.

    What does a model inventory need to include under SR 26-2?

    Every AI system in production or pilot, including third-party vendor tools with an embedded LLM, each assigned a risk tier based on the decision it influences.

    How long does an SR 26-2 gap assessment take?

    A typical roadmap runs about 90 days: inventory in the first three weeks, risk tiering and gap identification by week six, and closing the highest-risk gaps plus standing up ongoing governance by week twelve.

    Talk to the team that would do the work

    Bring your requirements to a working session with the person who'll actually deliver.

    Book a Discovery Call