Skip to content
AI Foundations for Bankers
0%

The Foundation Model Landscape

intermediate18 min readUpdated foundation-modelsgptclaudegeminillamaopen-weightcomparison
Jump to a section

From One Model to an Ecosystem

When the AI conversation began in earnest in late 2022, many people equated AI with a single product: ChatGPT. Today, the landscape has exploded. There are dozens of foundation models from major technology companies, open-source communities, and specialized AI labs -- each with distinct strengths, licensing terms, and deployment options.

For banking executives, this is actually good news. Competition drives innovation, reduces vendor lock-in risk, and creates options tailored to different regulatory and operational requirements. But navigating this landscape requires a structured approach.

KEY TERM

Foundation Model: A large AI model trained on broad data at scale that can be adapted (fine-tuned) to a wide range of downstream tasks. Unlike traditional models built for one specific purpose, a foundation model serves as the "foundation" upon which many applications can be built -- similar to how a core banking platform supports multiple product lines.

The Major Players

The foundation model market is led by a handful of providers, each bringing different philosophies and capabilities. Here is a snapshot of the model families most relevant to enterprise banking use cases. Each family comes in several tiers -- a top tier for the hardest work, a mid tier for everyday tasks, and a small, cheap tier for high volumes.

Model FamilyProviderKey StrengthBanking Use Case
GPTOpenAIBroad capability, large ecosystemDocument analysis, customer service automation
ClaudeAnthropicLong documents, careful reasoning, safety focusRegulatory document review, compliance analysis
GeminiGoogleMultimodal (text + image + video); an on-premises option on Google Distributed CloudCheck processing, document understanding
LlamaMetaOpen-weight, self-hosted deploymentInternal tools where data cannot leave the bank
CommandCohereEnterprise RAG with built-in citations; open-weight flagshipKnowledge base search, internal Q&A systems
MistralMistral AIEuropean-based, open-weight optionsEU banking operations, cross-border compliance

The strongest open-weight models -- models whose weights anyone can download and run -- now come largely from China: DeepSeek and Alibaba's Qwen family (some Qwen versions are open, some proprietary). For a bank, these raise a country-of-origin governance question as well as the usual capability question.

Suppliers also change strategy. In April 2026 Meta moved its frontier model to a closed line (Muse Spark), though it still releases smaller open-weight models. A supplier's change of direction is itself a vendor risk.

Tip

As of October 2026 -- current flagships. OpenAI: GPT-6 Astra (top), GPT-6.1 Sol (mid), GPT-6 Luna (low cost). Anthropic: Claude Fable 5.1 (top), Opus 5.5, Sonnet 5.5, Haiku 4.5. Google: Gemini 3.x (3.8 Flash is generally available; 3.1 Pro is still in preview). Cohere: Command A and Command A+. Mistral: Large 3 and Medium 3.5. Version numbers change every few months.

Each of these model families represents billions of dollars in training investment and distinct architectural choices. Selecting the right one -- or the right combination -- for your institution is a strategic decision, not a technical one.

BANKING ANALOGY

Choosing a foundation model is remarkably similar to your vendor evaluation process for a core banking system. You would never select a core platform based solely on a features checklist. You evaluate total cost of ownership, vendor stability, regulatory compliance capabilities, integration with existing infrastructure, and long-term strategic alignment. The same framework applies to foundation models. The "best" model in a benchmark may not be the best model for your institution.

Understanding the Differences

Capability and Quality

Not all foundation models are created equal. Performance varies significantly across different task types:

  • Reasoning and analysis: The top tiers of the GPT, Claude and Gemini families lead on complex reasoning tasks -- the kind of multi-step analysis required for credit risk assessment or regulatory interpretation
  • Long document processing: The current top models from OpenAI and Anthropic, and DeepSeek's latest open-weight model, all accept roughly 1 million tokens -- thousands of pages at once. A large context window is now table stakes
  • Code generation: The leading families all generate and explain code well, useful for data teams building analytics pipelines
  • Multilingual capability: Gemini and Mistral show strong performance across European and Asian languages, relevant for global banking operations
  • Retrieval-augmented tasks: Cohere's Command family is designed for RAG workflows, with built-in citations, making it a strong choice for internal knowledge management

Speed and Cost

There is a direct trade-off between model capability and operational cost, and the gap has widened: OpenAI's and Anthropic's top models list at $10 per million input tokens and $50 per million output tokens, while OpenAI's low-cost tier lists at $0.10 and $0.50 -- roughly a 100-fold difference. For many banking use cases, a less expensive model is entirely sufficient.

Consider the use case hierarchy:

  • High-stakes analysis (regulatory interpretation, risk assessment): Use the most capable model available. The cost per query is trivial compared to the value of accuracy
  • Document processing at scale (summarizing thousands of customer communications): Use a mid-tier model. You need good quality at reasonable cost
  • Simple classification tasks (routing customer inquiries, tagging documents): Use a smaller, faster model. Speed and cost matter more than peak capability

Data Privacy and Deployment Options

This is where banking requirements diverge sharply from other industries. When a retail company uses an LLM to generate marketing copy, data privacy is a minor concern. When a bank processes loan applications, customer financial records, or trading strategies through an LLM, data residency and privacy become paramount.

The deployment spectrum looks like this:

  • Provider cloud API (lowest control): Your data is sent to the provider's servers for processing. Enterprise agreements typically include data handling provisions, but the data does leave your perimeter
  • Your own cloud tenancy (moderate control): The model runs through the hyperscaler you already use -- AWS, Microsoft Azure or Google Cloud -- inside your own account, pinned to the regions you choose. Claude and GPT are both available this way
  • On-premises deployment (highest control): The model runs entirely within your infrastructure. This used to mean open-weight models only (such as Llama, Mistral or Command A+). Google now also offers Gemini air-gapped on Google Distributed Cloud, so a closed model from a frontier lab can run in the bank's own data center -- confirm which Gemini versions the on-premises option supports before planning around it. Maximum control, but more operational burden

Tip

For most banking institutions, the practical approach is a tiered strategy: match the deployment to the sensitivity of the data, not to a label like "open" or "closed." Public-document work can use a provider API under an enterprise agreement. Customer data can go to a closed model inside your own cloud tenancy with region pinning, or to an open-weight model on your own hardware.

Open-Weight vs. Closed Models

("Open-weight" is more accurate than "open-source": you get the trained model, usually not its training data.) The open-weight vs. closed debate is one of the most consequential decisions in your AI strategy. Here is what each approach offers:

Closed Models (GPT, Claude, Gemini)

Advantages:

  • Highest capability on complex tasks
  • Managed infrastructure -- no operational burden on your teams
  • Regular updates and improvements from the provider
  • Enterprise support agreements available

Disadvantages:

  • Usually means relying on cloud infrastructure (yours or the provider's); on-premises options are rare
  • Limited transparency into how the model works
  • Vendor lock-in risk -- switching costs increase over time
  • Pricing and model availability can change with limited notice

Open-Weight Models (Llama, Mistral, Command A+, DeepSeek, Qwen, NVIDIA Nemotron)

Advantages:

  • Full control over data -- nothing needs to leave your infrastructure
  • Ability to fine-tune on your own data (banking-specific terminology, institutional knowledge)
  • No per-query costs (after infrastructure investment)
  • No vendor dependency for the model itself

Disadvantages:

  • Requires significant ML engineering talent to deploy and maintain
  • Usually trails the leading closed models -- by months, not years
  • You bear the full operational burden (monitoring, scaling, updates)
  • Fine-tuning requires expertise and computing resources

The Hybrid Approach

Most sophisticated banking institutions are adopting a hybrid approach: closed models for high-complexity tasks, and open-weight models where the bank wants full control, fine-tuning or the lowest unit cost. This pragmatic middle ground maximizes capability while respecting the regulatory constraints that define banking.

Selection Criteria for Banking

When evaluating foundation models for your institution, structure the evaluation around these dimensions:

Regulatory Compliance

  • Does the provider offer data processing agreements that satisfy OCC/Fed/FDIC requirements?
  • Where is data processed and stored? Can you guarantee data residency within required jurisdictions?
  • Does the provider support audit trails and logging sufficient for regulatory examination?
  • What happens to your data after processing -- is it used for model training?

Operational Risk

  • What is the provider's uptime SLA? Banking operations often require 99.9%+ availability
  • What happens during provider outages? Do you have failover options?
  • How does the provider handle model updates? Can you control when you migrate to new model versions?
  • Is there model versioning so you can reproduce prior outputs for regulatory purposes?

Total Cost of Ownership

  • What is the per-query cost at your projected volume?
  • What internal infrastructure and talent is required?
  • What are the integration costs with your existing technology stack?
  • What is the switching cost if you need to change providers?

Strategic Alignment

  • Does this model strategy support your 3-5 year AI roadmap?
  • Are you building internal AI capability or outsourcing it?
  • How does this decision affect your talent strategy -- do you need ML engineers?

The model landscape is evolving rapidly. Any specific model comparison will be dated within months. What endures is the evaluation framework. Build institutional capability in evaluating models, not just in using any single one.

Looking Ahead

The foundation model landscape will continue to evolve rapidly. Models that are state-of-the-art today may be surpassed within months. New entrants will emerge, existing providers will release increasingly capable versions, and some will change strategy entirely.

For banking executives, the strategic imperative is clear: develop institutional fluency in evaluating and deploying foundation models, build the governance frameworks to manage them responsibly, and create the flexibility to adopt new models as the landscape evolves. The banks that treat AI model selection as a one-time decision will find themselves locked into yesterday's technology. The banks that build adaptive capability will thrive.

KNOWLEDGE CHECK

A bank needs to process customer loan applications through an LLM and wants maximum control over where the data goes. Which deployment approach best addresses data residency requirements?

What is the primary reason banking institutions are adopting a hybrid model strategy rather than selecting a single foundation model?

When evaluating foundation models, which factor most distinguishes banking's evaluation criteria from those of other industries?