Financial Services Governance
Model Risk Management for AI: From SR 11-7 to SR 26-2
How the shift from SR 11-7 to SR 26-2 changes model validation, inventory, and governance for generative and agentic AI.
context
What SR 26-2 Changes
Why Generative AI and LLMs Broke SR 11-7
SR 11-7, the Federal Reserve's 2011 model risk management guidance, was written for statistical and machine learning models with stable inputs and measurable accuracy. A large language model or an agentic AI system doesn't backtest the same way: its behavior can vary by prompt, its outputs are generative rather than a single predicted value, and a single deployment can span dozens of use cases rather than one model, one purpose.
Side-by-Side: What SR 11-7 Assumed vs. What SR 26-2 Addresses
SR 11-7 assumed a model has a fixed purpose, a measurable output, and a backtestable accuracy metric. SR 26-2 extends that framework to cover generative and agentic systems where validation has to focus on output quality sampling, guardrail testing, and use-case-specific risk tiering rather than a single accuracy number.
framework
Three Lines of Defense, Reinterpreted for AI
The first line (the business unit deploying the model), second line (independent risk and compliance oversight), and third line (internal audit) structure still applies, but AI adds new questions to each line: who owns a prompt-engineering change to a production LLM, who validates an agent's tool-use permissions, and who audits a vendor-supplied foundation model your team doesn't control the training data for.
validation
Model Validation for LLMs and Agentic AI
Why Backtesting Fails on Generative Models
A generative model doesn't produce one measurable prediction to check against ground truth. Validation instead needs structured output sampling against a rubric, red-teaming for guardrail failures, and ongoing drift monitoring on output quality rather than a single input-output accuracy check.
inventory
Model Inventory and Risk Tiering
Every AI system in production, including third-party vendor tools with an embedded LLM, needs to be in a model inventory with a risk tier assigned based on the decision it influences: a customer-facing chatbot answering general questions sits in a different tier than an agent that can initiate a wire transfer or approve a loan.
governance
Governance, Roles, and Board Reporting
SR 26-2 expects board and senior-management reporting on AI model risk with the same rigor as traditional model risk, meaning a bank or asset manager needs a defined owner for AI governance, not an informal arrangement where whoever deployed the model also validates it.
roadmap
A 90-Day Compliance Roadmap
A typical path: weeks one through three, inventory every AI system currently in production or pilot, including vendor tools with embedded AI. Weeks four through six, risk-tier each one and identify validation gaps against SR 26-2 expectations. Weeks seven through twelve, close the highest-risk gaps first and stand up the ongoing governance structure to keep the inventory current.
engagement
How Kriv AI Implements This
A regulated financial services client engaged Kriv AI to build an AI model inventory and validation framework ahead of an SR 26-2 gap assessment, the kind of anonymized engagement pattern we bring to a scoped model risk management project: inventory, risk tiering, validation framework design, and an ongoing governance structure your board can actually report against.
Straight answers
Frequently asked questions about Model Risk Management for AI: From SR 11-7 to SR 26-2
What is SR 26-2?
SR 26-2 is the Federal Reserve's updated model risk management guidance, extending the SR 11-7 framework to explicitly address generative AI, large language models, and agentic AI systems that don't validate the same way traditional statistical models do.
Does SR 11-7 still apply, or has it been replaced?
SR 26-2 builds on SR 11-7's three-lines-of-defense structure rather than discarding it, extending the validation and governance expectations to cover generative and agentic AI specifically.
Does SR 11-7 apply to large language models?
The original SR 11-7 guidance predates LLMs and doesn't address them explicitly, which is exactly the gap SR 26-2 is designed to close with generative-AI-specific validation expectations.
How do you validate an LLM if you can't backtest it like a traditional model?
Through structured output sampling against a rubric, red-teaming for guardrail failures, and ongoing drift monitoring on output quality, rather than a single input-output accuracy metric.
What does a model inventory need to include under SR 26-2?
Every AI system in production or pilot, including third-party vendor tools with an embedded LLM, each assigned a risk tier based on the decision it influences.
How long does an SR 26-2 gap assessment take?
A typical roadmap runs about 90 days: inventory in the first three weeks, risk tiering and gap identification by week six, and closing the highest-risk gaps plus standing up ongoing governance by week twelve.
Talk to the team that would do the work
Bring your requirements to a working session with the person who'll actually deliver.
Book a Discovery Call