Model Risk Management (MRM) for LLMs
Jump to a section
Why LLMs Demand Special MRM Attention
Banking institutions have decades of experience managing model riskModel Risk ManagementSupervisory guidance and bank practices governing how models are inventoried, validated, monitored, and controlled. Since April 2026 the interagency guidance is SR 26-2 (with OCC Bulletin 2026-13 and FDIC FIL-15-2026), which replaced SR 11-7 and OCC Bulletin 2011-12. Generative and agentic AI are explicitly outside its scope, but examiners can still act on unsafe or unsound practices or violations of law, and most banks continue to apply model-risk disciplines to their AI systems.See glossary. Credit scoring models, market risk models, anti-money laundering models -- these are mature disciplines with well-established validation methodologies, governance structures, and supervisory expectations. Your institution almost certainly has a Model Risk Management framework built on SR 11-7, the Federal Reserve's 2011 guidance that defined bank MRM for a decade and a half.
Large Language ModelsLarge Language Model (LLM)A neural network trained on vast amounts of text data that can understand and generate human language. LLMs power chatbots, document analysis, code generation, and many enterprise AI applications.See glossary break nearly every assumption that traditional MRM frameworks are built on. And in 2026 the US guidance itself changed -- in a way that puts more of the responsibility for governing LLMs on the bank, not less.
KEY TERM
Model Risk Management (MRM): The discipline of identifying, measuring, monitoring, and controlling the risk that arises from the use of models in business decisions. In the United States, bank MRM is now guided by SR 26-2 (2026), which replaced SR 11-7; the UK, Canada and other jurisdictions have their own frameworks. MRM encompasses model development, validation, ongoing monitoring, and governance.
From SR 11-7 to SR 26-2
In April 2026 the Federal Reserve, OCC and FDIC replaced SR 11-7 (and OCC Bulletin 2011-12) with updated model risk management guidance -- SR 26-2, OCC Bulletin 2026-13 and FDIC FIL-15-2026. The new guidance is aimed mainly at banks with over $30 billion in assets, and it explicitly places generative and agentic AI outside its scope while the agencies gather input on how banks use AI (the OCC said a request for information is coming). That is not a free pass: examiners can still act on unsafe or unsound practices or violations of law, and most banks continue to apply model-risk disciplines -- inventory, validation, monitoring, documentation -- to their AI systems.
What Changed
- A narrower definition of "model." A model is now a "complex quantitative method" that applies statistical, economic, or financial theories. The word "mathematical" is gone, and spreadsheets and deterministic rule-based processes are expressly excluded.
- Tailored to size. The guidance is "most relevant to" banking organizations with more than $30 billion in assets. Smaller banks are generally outside it unless they have significant model risk.
- Principles, not a checklist. Validation is risk-based rather than on a fixed cycle. Independence is framed as "effective challenge" by people with "sufficient independence." A model inventory is described as "common industry practice." Vendor and third-party models still get their own section.
- Generative and agentic AI are carved out. A footnote calls these models "novel and rapidly evolving" and places them outside the guidance; the bank's own risk management and governance practices should guide the controls. The same footnote confirms the principles do apply to non-generative, non-agentic AI -- so a machine-learning credit or fraud model is still squarely in scope.
- Older companion guidance withdrawn. The OCC also withdrew its separate bulletins on BSA/AML model risk and on credit scoring models.
The Footnote Every Executive Should Read
The guidance says it "does not set forth enforceable standards" and that not following it "will not result in supervisory criticism." Read that sentence alone and you might conclude MRM is now optional. Footnote 1 is the other half: supervisory action can still follow from violations of law or from unsafe or unsound practices. An LLM that gives customers wrong fee information, leaks data, or shapes credit decisions without controls is not protected by the scope carve-out. An examiner simply reaches it through safety-and-soundness, consumer protection or fair lending authority instead of the model risk guidance.
What It Means for Your AI Governance Now
- Your framework is the standard. For generative AI there is currently no US supervisory MRM standard. Whatever your bank writes down -- tiering, testing, approval, monitoring -- is what you will be measured against. Write it as if an examiner will read it.
- Do not dismantle what works. Inventory, independent challenge, validation and monitoring remain the best evidence that an AI system is safe and sound. Most banks are keeping them.
- Size matters, but it is not an exemption from risk. A bank below $30 billion is mostly outside the guidance, but still answers for unsafe practices and consumer harm.
- Watch for the request for information. The OCC said the agencies plan to ask the industry how banks use AI, including generative and agentic AI. Expect US expectations for LLMs to become more specific after that, and plan for your framework to need another pass.
BANKING ANALOGY
SR 26-2 treats generative AI the way a credit policy treats a new product it has not yet priced: it does not say "anything goes," it says "this is not covered by the standard grid -- the credit committee owns the decision." The bank still has to underwrite the risk; it just cannot point to the grid as its justification.
Why LLMs Strain Traditional MRM
The core disciplines -- sound development, effective challenge, governance -- carry forward. Applying them to LLMs is much harder. Here is why.
The Opacity Problem
Traditional credit models produce a score, and you can trace how each input variable contributed to it. LLMs are fundamentally opaque. With billions of parameters, there is no practical way to trace why the model generated a specific output. The explainability that validators and examiners are used to becomes extraordinarily difficult to provide.
The Non-Determinism Problem
Run the same input through a traditional credit model ten times, and you get the same output ten times. Run the same prompt through an LLM ten times, and you may get ten different responses -- all potentially valid, but none identical. This challenges every testing and validation method built for traditional models.
The Scope Problem
A traditional model has a clearly defined scope: it scores credit risk, or it detects fraud. An LLM deployed as a general-purpose assistant could be used for tasks no one anticipated -- customer communications one moment, regulatory analysis the next. Defining the "intended use" of an LLM, a core MRM discipline, is far more complex.
The Training Data Problem
For traditional models, you control the training data. LLMs are trained on internet-scale data that you did not curate, cannot fully audit, and may contain biases you cannot identify. If your institution fine-tunesFine-TuningThe process of further training a pre-trained model on a specific dataset to specialize its behavior for a particular domain or task, such as banking compliance language.See glossary a model on proprietary data, you add another layer of complexity.
Warning
The scope carve-out for generative AI is not a holiday from risk management. The agencies' 2026 model risk guidance leaves generative AI to each bank's own governance, and examiners can still cite unsafe or unsound practices. Institutions that deploy LLMs without a documented framework face regulatory, legal, and reputational risk. Do not treat LLM deployment as a technology project that can proceed outside your risk management framework.
The Three Lines of Defense for LLMs
Banking institutions typically organize risk management around three lines of defense. Here is how each line must adapt for LLMs:
First Line: Business Units and Model Owners
- Use case documentation: Clearly defining how the LLM is being used, what decisions it informs, and what customer-facing outputs it produces
- Input quality controls: Establishing guardrailsGuardrailsSafety mechanisms that constrain AI model outputs, and the actions AI agents can take, to prevent harmful, off-topic, non-compliant, or unauthorized results. Critical in banking for regulatory adherence and brand safety.See glossary on what data can be sent to the LLM and what prompts are permitted
- Output review processes: Human review for high-risk outputs (anything customer-facing, anything regulatory, anything involving lending decisions)
- Incident reporting: Flagging hallucinationsHallucinationWhen an AI model generates plausible-sounding but factually incorrect information. A critical risk in banking where inaccurate outputs could lead to regulatory violations or financial losses.See glossary, inappropriate outputs, or unexpected behavior to the second line
Second Line: Risk Management and Compliance
- LLM model inventory: Maintaining a comprehensive inventory of all LLM deployments, including shadow IT usage (employees using consumer AI tools for work purposes)
- Risk tiering: Classifying LLM use cases by risk level -- a customer service chatbot has different risk implications than an LLM-assisted credit decisioning tool
- Validation methodology: Developing validation approaches appropriate for LLMs (more on this below)
- Policy development: Establishing acceptable use policies, data handling requirements, and governance standards
Third Line: Internal Audit
- Assess framework effectiveness: Evaluate whether the institution's LLM governance is adequate for the risk posed
- Test controls independently: Verify that first and second line controls are operating as designed
- Evaluate regulatory compliance: Confirm that LLM deployments satisfy applicable legal and regulatory requirements
- Audit trail review: Verify that model inputs, outputs, and decisions are logged and retrievable
Validation Challenges for LLMs
Benchmark Testing
Instead of validating against a single quantitative outcome (does the model correctly predict default?), LLM validation requires testing across a broad range of scenarios with qualitative assessment of output quality. Leading institutions are building "test suites" -- curated sets of prompts with expert-assessed reference answers -- specific to each use case.
Bias and Fairness Testing
Fair lending law prohibits discrimination on protected characteristics. For LLMs, you need to test whether the model produces systematically different language, tone, or recommendations when the input varies only by protected characteristics. This is an emerging discipline with limited precedent.
Ongoing Monitoring
LLMs can behave unpredictably when:
- The model provider updates the underlying model (often without advance notice)
- Users discover novel prompting techniques that circumvent guardrails
- The model encounters edge cases not covered by initial testing
Effective monitoring requires continuous evaluation rather than a fixed calendar -- consistent with the risk-based approach the 2026 guidance takes for models generally. Many institutions sample LLM outputs automatically and flag anomalies for human review.
Red Team Testing
Red team testing means deliberately trying to make the LLM produce harmful, inaccurate, or non-compliant outputs. For banking applications, red teams should attempt to:
- Elicit outputs that violate fair lending requirements
- Generate plausible but inaccurate financial information
- Circumvent data handling restrictions
- Produce outputs that could constitute unauthorized investment advice
Building an LLM MRM Framework
Risk Tiering
Classify every LLM use case into risk tiers:
- Tier 1 (Critical): LLM outputs directly influence lending decisions, customer pricing, regulatory submissions, or financial reporting. Full validation required. Human review of every output mandatory.
- Tier 2 (Significant): LLM outputs are customer-facing or inform material business decisions. Validation required. Sampling-based human review.
- Tier 3 (Standard): LLM used for internal productivity (drafting, summarization, research). Acceptable use policy compliance. User training required.
Governance Structure
- AI/Model Risk Committee: Cross-functional body with authority over LLM deployment decisions
- Model owners: Named individuals accountable for each LLM use case
- Validation team: Independent team (or external firm) with LLM-specific expertise
- Executive sponsorship: Senior leader accountable to the board for AI risk management
Documentation Requirements
At minimum, document for each LLM deployment: the use case and intended scope; why this model was selected; data handling (what goes in, where it goes); prompt design and guardrails; testing and validation results; the ongoing monitoring plan; and incident response procedures. With no US supervisory template for generative AI, this documentation is your evidence of sound governance.
Tip
Do not wait for a perfect framework before deploying LLMs. Deploy Tier 3 use cases (internal productivity) under your existing acceptable use policies while you build the full governance framework for Tier 1 and Tier 2. This lets your institution gain experience with the technology while managing risk appropriately.
The Regulatory Trajectory
Dated facts to anchor your planning (as of October 2026):
- United States: SR 26-2 leaves generative and agentic AI to bank governance, and the OCC has said an AI request for information is coming. A December 2025 executive order ("Ensuring a National Policy Framework for Artificial Intelligence") targets "onerous" state AI laws, and the Department of Justice set up an AI Litigation Task Force in January 2026.
- States: Colorado's original AI Act never took effect; it was repealed and replaced by a new law that takes effect in January 2027 (see Responsible AI & Fairness).
- EU AI Act: Obligations for general-purpose AI models have applied since August 2025, with Commission enforcement from August 2026. The Digital Omnibus on AI (in force since July 2026) moved the high-risk deadline -- which covers credit scoring -- to December 2027.
- UK: The PRA's SS1/23 model risk principles (applying since May 2024) cover AI and machine-learning models -- but only for banks, building societies and investment firms with internal-model approval for regulatory capital.
- Canada: OSFI's Guideline E-23 takes effect in May 2027 for all federally regulated financial institutions and covers AI and machine-learning models.
- NIST: The AI Risk Management Framework is being revised as part of the White House AI Action Plan; its Generative AI Profile (July 2024) is unchanged.
Notice the contrast: Canada and, for internal-model firms, the UK are now more explicit about AI in model risk than US guidance is. Banks operating across borders will find the strictest home supervisor sets the practical bar.
Moving Forward
Model Risk Management for LLMs is not a solved problem -- it is an evolving discipline. The principles of sound development, effective challenge and governance carry forward in the 2026 guidance. For LLMs, applying them is now the bank's own choice and responsibility, not a supervisory checklist. The banks that build on their existing MRM expertise, rather than treating the carve-out as permission to stop, will be best placed when US expectations for AI firm up.
KNOWLEDGE CHECK
Why do LLMs present a fundamentally different challenge for model validation compared to traditional credit scoring models?
Under the three lines of defense model, which responsibility belongs to the SECOND line (Risk Management and Compliance) when governing LLM deployments?
A bank wants to deploy an LLM-powered tool that assists credit analysts with loan recommendations. Under the risk tiering framework described, how should this use case be classified?