Model Selection Framework for Banking
Jump to a section
From Model Awareness to Model Strategy
You now understand the major foundation modelsFoundation ModelA large AI model trained on broad data that can be adapted to many tasks. Examples include OpenAI's GPT, Anthropic's Claude, Google's Gemini and Meta's Llama families. Banks evaluate these for capabilities, safety, and regulatory fit.See glossary -- their strengths, trade-offs, and deployment options. But understanding individual models is not the same as having a model strategy. This unit provides the framework your institution needs to make structured, defensible model selection decisions for each banking use case.
The framework addresses four dimensions: Capability (can the model do the job?), Cost (at what total cost of ownership?), Compliance (does it satisfy regulatory requirements?), and Control (what level of oversight and customization do you retain?). Every model selection decision requires balancing these four dimensions against the specific requirements of the use case.
The Four-Dimension Framework
Dimension 1: Capability
Not every banking task requires the most powerful model. Match model capability to task complexity:
| Task Complexity | Example Banking Tasks | Model Tier |
|---|---|---|
| Low | Document classification, email routing, FAQ response | Small/fast models (Claude Haiku, OpenAI's low-cost GPT tier, Gemini Flash-Lite, small open-weight models such as Ministral or Nemotron Nano) |
| Medium | Document summarization, draft generation, data extraction | Mid-tier models (Claude Sonnet, OpenAI's mid GPT tier, Gemini Flash, Mistral Medium) |
| High | Regulatory interpretation, complex risk analysis, multi-step reasoning | Top-tier models (Claude Opus or Fable, OpenAI's top GPT tier; Google's top Gemini Pro model was still in preview in October 2026) |
Over-provisioning capability wastes money. A simple email classification task does not need a top-tier model -- a model 10x cheaper will perform equally well. Under-provisioning capability creates risk. A regulatory interpretation task needs the strongest available reasoning.
BANKING ANALOGY
This is the same principle you apply to staffing decisions. You do not assign a Managing Director to process routine wire transfers, and you do not assign a junior analyst to structure a complex syndicated loan. Each task has an appropriate expertise level. AI model selection follows the same logic -- match the model's capability to the task's complexity, and deploy your most capable (and expensive) resources only where they add the most value.
Dimension 2: Cost
Total cost of ownership varies dramatically based on deployment approach:
Per-query pricing (cloud APIs):
- Best for: variable workloads, lower volumes, rapid prototyping
- Cost driver: token volume. Price ranges from roughly $0.10 per million tokens (cheapest input) to $50 per million tokens (top-tier output)
- Cost levers: batch processing (50% off for work that can wait) and caching repeated content, such as a policy manual (about 90% or more off)
- Watch out for: costs scaling linearly with usage -- a successful deployment can become expensive quickly
Managed infrastructure (Together AI, Amazon Bedrock, Microsoft Foundry, Google's Gemini Enterprise Agent Platform -- formerly Vertex AI):
- Best for: steady-state workloads, moderate volumes, when you want model variety without infrastructure management
- Cost driver: provisioned capacity. You pay for reserved compute regardless of utilization
- Watch out for: over-provisioning capacity during low-demand periods
- Bonus: platforms such as Bedrock host several vendors' models side by side, which makes switching easier
Self-hosted (on-premises or dedicated cloud):
- Best for: high-volume workloads, maximum data control, predictable costs at scale
- Cost driver: hardware, operations, talent. Fixed costs regardless of query volume
- Watch out for: underestimating the operational burden and talent requirements
KEY TERM
Total Cost of Ownership (TCO): The complete cost of deploying and operating an AI model, including API fees or hardware costs, engineering time for integration and maintenance, monitoring and governance overhead, and the opportunity cost of model management versus other IT priorities. For banking AI decisions, TCO analysis should span at least a 3-year horizon.
Dimension 3: Compliance
For banking, compliance is not optional -- it is a gating criterion that can eliminate model options entirely:
Data residency: Where is data processed? Can you guarantee geographic boundaries? Cloud APIs may route to servers in any region. VPC and on-premises deployments provide geographic control.
Audit trails: Can you log every prompt, response, and model version for regulatory examination? Enterprise API agreements typically include logging. Self-hosted deployments require building this capability.
Model risk managementModel Risk ManagementSupervisory guidance and bank practices governing how models are inventoried, validated, monitored, and controlled. Since April 2026 the interagency guidance is SR 26-2 (with OCC Bulletin 2026-13 and FDIC FIL-15-2026), which replaced SR 11-7 and OCC Bulletin 2011-12. Generative and agentic AI are explicitly outside its scope, but examiners can still act on unsafe or unsound practices or violations of law, and most banks continue to apply model-risk disciplines to their AI systems.See glossary: Can you document model selection rationale, validation testing, and ongoing monitoring? Models with published training data (some open-weight models publish it) provide more transparency for MRM documentation.
In April 2026 the Federal Reserve, OCC and FDIC replaced SR 11-7 (and OCC Bulletin 2011-12) with updated model risk management guidance — SR 26-2, OCC Bulletin 2026-13 and FDIC FIL-15-2026. The new guidance is aimed mainly at banks with over $30 billion in assets, and it explicitly places generative and agentic AI outside its scope while the agencies gather input on how banks use AI (the OCC said a request for information is coming). That is not a free pass: examiners can still act on unsafe or unsound practices or violations of law, and most banks continue to apply model-risk disciplines — inventory, validation, monitoring, documentation — to their AI systems.
In practice there is no US model-risk rulebook written for large language models, so your own governance judgement fills the gap -- and your selection framework is your evidence.
Data handling: What happens to your data after processing? Is it used for model training? Enterprise agreements should explicitly prohibit training on your data. Self-hosted models eliminate this concern entirely.
Tip
Build a compliance checklist specific to your institution's regulatory requirements and apply it consistently to every model evaluation. This transforms model selection from a subjective technical preference into a structured, auditable governance process -- exactly the evidence of sound AI risk management examiners look for.
Dimension 4: Control
Control encompasses customization, vendor dependency, and operational autonomy:
Customization: Can you fine-tuneFine-TuningThe process of further training a pre-trained model on a specific dataset to specialize its behavior for a particular domain or task, such as banking compliance language.See glossary the model on your data? Open-source models offer full fine-tuning. Some proprietary models offer limited fine-tuning. Others offer none.
Vendor dependency: What happens if the vendor changes pricing, deprecates your model version, or modifies data handling terms? Proprietary models create higher vendor dependency. Open-source models you host eliminate vendor dependency for the model itself (but create dependency on your own operations team).
Version control: Can you lock to a specific model version for reproducibility? Regulatory environments may require reproducing prior outputs. Version pinning is essential -- but pinned versions do get retired. Some vendors (Anthropic, for example) publish a minimum-availability date for each model version, and cloud resellers may set their own retirement calendars. When OpenAI ended API access to one of its older models in February 2026, any bank pinned to it had to migrate. Plan migrations before you need them.
Portability: How difficult is it to switch models if your current choice becomes suboptimal? Architectures that abstract the model behind an API interface make switching easier.
Model Comparison for Banking
| Model Family | Capability | Cost (relative) | Compliance Readiness | Control |
|---|---|---|---|---|
| Claude (Anthropic) | Top-tier reasoning, 1M context, strong safety | $$$ | Enterprise agreements, no data training; Bedrock, Google Cloud, Microsoft Foundry | Limited fine-tuning, cloud-only |
| GPT (OpenAI) | Top-tier general capability, multimodal, 1M context | $$$ | Enterprise agreements; Azure and Amazon Bedrock (Microsoft's cloud exclusivity ended April 2026) | Limited fine-tuning, cloud-only |
| Gemini (Google) | Top-tier multimodal capability | $$$ | Enterprise agreements; Google Cloud, plus an air-gapped on-premises option (Google Distributed Cloud; confirm which versions it supports) | Limited fine-tuning; air-gapped on-prem option |
| Command A / A+ (Cohere) | Best-in-class RAG with citations, multilingual | $$ | VPC/on-prem options, data residency controls | Full control with the open-weight Command A+ (Apache 2.0) |
| Open-weight leaders (Llama, Qwen, DeepSeek, Nemotron) | Strong, usually months behind the frontier | $ (self-hosted) | Full data control, on-premises capable; review licence and country of origin | Full fine-tuning, full control |
| Mistral (Mistral AI) | Strong open-weight, EU-based | $ (self-hosted) | GDPR alignment, on-premises capable | Full fine-tuning, full control |
Warning
This comparison reflects the landscape at publication time. Model capabilities, pricing, and compliance features evolve rapidly. The framework itself is the durable asset -- apply these four dimensions to any new model that enters the market. Do not treat this table as a permanent ranking.
Matching Use Cases to Models
Recommended Model Tiers by Banking Use Case
| Use Case | Data Sensitivity | Complexity | Recommended Approach |
|---|---|---|---|
| Regulatory document analysis | Medium | High | Top-tier proprietary model via enterprise API |
| Internal policy Q&A (RAG) | High | Medium | RAG-specialist model (VPC) or open-weight model (on-prem) with RAG pipeline |
| Customer email classification | Medium | Low | Small open-weight model on-premises |
| Credit memo draft generation | High | High | Top-tier proprietary via enterprise API, with human review |
| Market research summarization | Low | Medium | Mid-tier proprietary model via standard API |
| Code generation for data teams | Low | Medium | Mid-tier model via API |
| Customer complaint analysis | High | Medium | On-premises open-weight model with fine-tuning |
| Board presentation drafting | Low | High | Top-tier proprietary model via enterprise API |
Building Your Model Strategy
A mature banking AI strategy does not rely on a single model. It builds a tiered architecture:
Tier 1 -- Strategic analysis: Top-tier proprietary models for complex reasoning, regulatory interpretation, and high-stakes analysis. Low volume, high value per query.
Tier 2 -- Operational intelligence: Mid-tier or specialized models (a RAG specialist for knowledge search, a mid-tier general model for everything else) for steady-state business applications. Moderate volume, moderate value per query.
Tier 3 -- High-volume automation: Small, efficient models (often open-weight and fine-tuned) for classification, routing, extraction, and other high-volume tasks. High volume, lower value per query.
This tiered approach ensures you deploy the right capability at the right cost for each use case, while maintaining compliance and control where they matter most.
BANKING ANALOGY
This tiered model strategy mirrors how your institution manages its investment portfolio. You do not put all assets in a single instrument. You maintain a diversified portfolio: high-conviction, high-return positions (top-tier models for strategic analysis), core holdings for steady returns (mid-tier models for operational use), and efficient, low-cost positions for broad market exposure (small models for high-volume automation). The portfolio is managed holistically, with risk and return balanced across tiers.
Governance Integration
Model selection is not a one-time decision. It is an ongoing governance process:
- Initial evaluation: Apply the four-dimension framework to select models for each use case
- Validation testing: Test selected models against banking-specific benchmarks before deployment
- Ongoing monitoring: Track model performance, cost, and hallucinationHallucinationWhen an AI model generates plausible-sounding but factually incorrect information. A critical risk in banking where inaccurate outputs could lead to regulatory violations or financial losses.See glossary rates in production
- Periodic re-evaluation: Review model selections quarterly as new models and capabilities emerge
- Deprecation planning: Track each vendor's retirement dates and maintain migration paths so you can switch models without disrupting operations
Quick Recap
- The four-dimension framework evaluates models on Capability, Cost, Compliance, and Control -- balancing all four for each use case
- Match model capability to task complexity -- over-provisioning wastes money, under-provisioning creates risk
- Total cost of ownership varies dramatically between cloud APIs, managed infrastructure, and self-hosted deployments
- A mature banking model strategy uses a tiered approach: top-tier for strategic analysis, mid-tier for operations, efficient models for high-volume automation
- Model selection is an ongoing governance process, not a one-time decision -- integrate it with your Model Risk Management framework
- SR 26-2 (April 2026) places generative and agentic AI outside its scope, so your documented selection framework is your evidence of sound AI risk management
KNOWLEDGE CHECK
A bank needs to classify 2 million customer emails per month by topic. Using the four-dimension framework, which model approach is most appropriate?
Why does this framework recommend different models for regulatory document analysis versus customer email classification?
What is the primary risk of a banking institution relying on a single foundation model for all AI use cases?