Skip to content
AI Foundations for Bankers
0%

Hugging Face & Together.ai — Model Marketplaces

intermediate10 min readUpdated hugging-facetogether-aiopen-sourceopen-weightmodel-hubinference
Jump to a section

The Open-Weight Model Revolution

The foundation model landscape is not limited to proprietary models from OpenAI, Anthropic, and Google. A parallel revolution is happening in open AI, where powerful models are freely available for anyone to download, deploy, and modify. These are usually called "open-weight" models: you get the trained model itself, though usually not the data it was trained on. For banking institutions, this ecosystem offers something most proprietary models cannot: complete control over your AI infrastructure.

Two companies have become the gateways to this ecosystem. Hugging Face operates the world's largest repository of open AI models. Together AI provides the enterprise-grade infrastructure to run those models efficiently. Together, they represent a compelling alternative -- or complement -- to proprietary model vendors.

Hugging Face: The GitHub of AI

Hugging Face Hub is to AI models what GitHub is to source code: a central repository where researchers and companies publish, share, and collaborate on models. With more than 3 million public models available (August 2026), Hugging Face has become the default distribution platform for the open AI community.

What Hugging Face Offers

Model repository: Browse, evaluate, and download models ranging from compact 1-billion-parameter models to massive models with hundreds of billions of parameters. The most-used families today include Qwen (Alibaba), DeepSeek, Llama (Meta), Mistral, Gemma (Google), NVIDIA Nemotron, OpenAI's gpt-oss and Cohere's Command A+ -- plus thousands of fine-tuned variants.

Model cards: Standardized documentation for each model -- training data, performance benchmarks, intended use cases, known limitations. This transparency is valuable for banking model risk management, where you need to document model provenance and characteristics.

Transformers library: An open-source software library that provides a unified interface for loading and running models. Your data science team can switch between model architectures with minimal code changes.

Spaces: Interactive demo environments where you can test models before committing to deployment. Evaluate a model's performance on your specific use cases before investing in infrastructure.

BANKING ANALOGY

Think of Hugging Face like the secondary market for banking assets. Just as your institution evaluates and acquires loan portfolios or securities from the secondary market -- applying your own risk criteria, due diligence, and management practices -- Hugging Face lets you evaluate and acquire AI models from the open market. You can inspect the model's "portfolio" (training data), review its "performance history" (benchmarks), and deploy it under your own risk management framework. The model is the asset; Hugging Face is the marketplace.

A Change of Ownership

In September 2026 NVIDIA -- the dominant supplier of AI chips -- agreed to acquire Hugging Face, for about $11.9 billion plus up to $1 billion in retention equity. The deal is expected to close in the first half of 2027, subject to regulatory approval, and NVIDIA has committed to keeping the platform open. For banks, it is a concentration-risk point worth noting: the neutral marketplace for open models is set to belong to the company that also supplies most of the hardware they run on.

Together AI: Enterprise Inference Infrastructure

Downloading an open-weight model is free. Running it efficiently at enterprise scale is not. Together AI, which now describes itself as the "AI Native Cloud," bridges this gap by providing optimized inference infrastructure for open models. It raised $800 million at an $8.3 billion valuation in July 2026, and in August 2026 signed a multi-year agreement with IBM to run open-model inference on IBM Cloud -- a familiar enterprise route in for many banks.

The Inference Challenge

Running large AI models requires specialized hardware -- typically NVIDIA GPUs with sufficient memory and compute power. A 70-billion-parameter model might require 4-8 high-end GPUs just to load into memory. Managing this infrastructure -- provisioning, scaling, monitoring, optimizing -- requires specialized ML engineering talent.

Together AI handles this complexity by offering:

  • Managed inference endpoints: Run any major open-weight model through a simple API call, without managing GPU infrastructure
  • Fine-tuning platform: Fine-tune open-weight models on your own data, creating banking-specific model variants
  • Cost optimization: Together AI says its optimized serving infrastructure delivers lower inference costs than running models on raw cloud GPU instances -- a vendor claim to test against your own workloads
  • Model variety: Switch between different open-weight model families through the same API interface

KEY TERM

Model Marketplace: A platform that aggregates AI models from multiple providers and researchers, providing standardized access, documentation, and (in some cases) optimized deployment infrastructure. Hugging Face is the marketplace; Together AI is one of the infrastructure providers that makes marketplace models production-ready.

Banking Use Cases for Open-Weight Models

Open-weight models offer distinct advantages for specific banking scenarios:

On-Premises Deployment for Sensitive Data

When processing customer financial data, proprietary trading strategies, or regulatory examination materials, many banks require that no data leave their infrastructure. Open-weight models can be deployed entirely on-premises -- in your own data center, on your own hardware, with no external API calls.

Cost Control at Scale

Proprietary model APIs charge per token. For high-volume use cases -- processing millions of customer communications, classifying thousands of transactions daily, or summarizing weeks of market data -- these costs add up. Open-weight models deployed on your own infrastructure have fixed costs (hardware and operations) regardless of volume, which may be more economical at scale.

Fine-Tuning for Banking-Specific Tasks

Open-weight models can be fine-tuned on your institution's data -- regulatory filings, credit memos, compliance correspondence -- to create models that deeply understand banking terminology and conventions. This level of customization is typically not available with proprietary models.

Vendor Independence

Relying on a single proprietary model provider creates concentration risk. If that provider changes pricing, deprecates a model version, or modifies their data handling policies, your institution is affected. Open-weight models reduce this dependency -- you hold the model weights and can run them indefinitely. But the supplier of the next version can still change course, and the hubs you download from are part of your software supply chain, so the dependency shrinks rather than disappears.

Tip

For most banking institutions, the optimal strategy is not "open OR proprietary" but "open AND proprietary." Use the top tiers of the major proprietary families for complex reasoning tasks where capability is paramount. Use open-weight models for high-volume tasks, sensitive data processing, and cost optimization. This hybrid approach maximizes both capability and control.

Evaluating Open-Weight Models for Banking

When evaluating open-weight models, apply the same rigor you would to any model under your institution's Model Risk Management framework:

Evaluation DimensionKey Questions
Model provenanceWho created the model, and in which country? What data was it trained on? Are there known biases?
LicensingDoes the license permit commercial use? Are there restrictions on financial services applications, or conditions tied to company size or revenue?
PerformanceHow does it perform on banking-relevant benchmarks? (Regulatory text comprehension, financial reasoning, compliance classification)
Operational burdenWhat hardware is required? What ML engineering talent is needed for deployment and maintenance?
Community supportIs the model actively maintained? Are security vulnerabilities being patched? Is the hub you download from treated as part of your software supply chain?

Warning

Not all open model licenses are the same, and "open" now covers a wide range. Some models use permissive licenses -- Apache 2.0 (for example Cohere Command A+) -- that allow broad commercial use. Others use modified permissive licenses (Mistral Medium 3.5 uses a modified MIT licence), and some use custom community licences, such as Meta's Llama licence, that add conditions which can depend on the size or revenue of the company using them. Always review the specific license terms with your legal team before deploying an open-weight model in a banking context.

The Practical Path Forward

For a banking institution exploring open-weight AI, a practical approach is:

  1. Evaluate on Hugging Face Spaces: Test promising models against your actual use cases before any infrastructure investment
  2. Prototype with Together AI: Build a proof-of-concept using Together AI's managed inference -- no GPU procurement needed
  3. Fine-tune for your domain: Use Together AI's fine-tuning platform to create a banking-specific model variant
  4. Decide on deployment: Based on results, choose between continued managed inference (Together AI) or bringing the model in-house on your own infrastructure

This graduated approach lets your institution build confidence and capability incrementally, without committing to large infrastructure investments upfront.

Quick Recap

  • Hugging Face is the world's largest repository of open AI models, offering 3 million+ public models with standardized documentation -- and NVIDIA has agreed to acquire it
  • Together AI provides enterprise inference infrastructure that makes running open-weight models practical without managing GPU hardware
  • Open-weight models enable on-premises deployment for sensitive data, cost control at scale, fine-tuning for banking-specific tasks, and greater vendor independence
  • The optimal banking strategy combines proprietary models for complex reasoning with open-weight models for high-volume and sensitive-data use cases
  • Evaluate open-weight models -- including their licence and country of origin -- with the same rigor as any model under your Model Risk Management framework

KNOWLEDGE CHECK

What is the primary strategic advantage of open-weight AI models for banking institutions compared to proprietary models?

A bank processes 5 million customer emails monthly through an AI classification system. Why might open-weight models be more cost-effective than proprietary APIs for this use case?

Why does the evaluation of open-weight models for banking require reviewing model licensing terms?