What Is a Large Language Model (LLM)?
Jump to a section
The Technology Reshaping Financial Services
Every decade or so, a technology arrives that fundamentally changes how banks operate. In the 1990s, it was online banking. In the 2010s, it was mobile-first platforms. Today, Large Language ModelsLarge Language Model (LLM)A neural network trained on vast amounts of text data that can understand and generate human language. LLMs power chatbots, document analysis, code generation, and many enterprise AI applications.See glossary represent the next inflection point -- and understanding them is no longer optional for banking leaders.
But here is the good news: you do not need a computer science degree to grasp what LLMs are and why they matter. You need the same analytical framework you already use when evaluating any major technology investment.
KEY TERM
Large Language Model (LLM): A type of artificial intelligence (foundation modelFoundation ModelA large AI model trained on broad data that can be adapted to many tasks. Examples include OpenAI's GPT, Anthropic's Claude, Google's Gemini and Meta's Llama families. Banks evaluate these for capabilities, safety, and regulatory fit.See glossary) trained on massive volumes of text data -- billions of documents, books, articles, and web pages -- that can understand context, generate human-like text, summarize information, translate languages, and answer questions. Today's leading models are also multimodalMultimodalDescribes models or embeddings that handle images, scanned documents, and audio as well as text -- for example, reading a scanned loan package or a cheque image.See glossary: they can read images and scanned documents as well as text. Unlike traditional software that follows explicit rules, an LLM learns patterns from data and applies them to new situations.
How LLMs Actually Work
At the most fundamental level, an LLM is a prediction engine. Given a sequence of words, it predicts what comes next. That sounds simple, but when you train this prediction engine on trillions of words and give it billions of internal parameters to adjust, something remarkable happens: the model develops an apparent understanding of language, logic, and even domain-specific knowledge.
Here is a simplified view of the process:
- Training: The model reads enormous amounts of text -- regulatory filings, news articles, scientific papers, books, code repositories -- and learns statistical patterns about how language works
- Parameters: During training, the model adjusts billions of internal weights (think of these as dials that get fine-tuned). Leading models have hundreds of billions to trillions of parameters, but vendors rarely disclose the count, and size alone no longer predicts quality
- InferenceInferenceThe process of running a trained model to generate predictions or outputs from new input data. Inference cost, latency, and throughput are key factors in enterprise AI deployment.See glossary: When you ask the model a question, it uses those learned patterns to generate a response, one word (technically, one "tokenTokensThe basic units of text that LLMs process. A token is typically one-half to three-quarters of an English word, depending on the model. Token counts determine cost, speed, and context window limits for every API call.See glossary") at a time
- Context windowContext WindowThe maximum amount of text (measured in tokens) a model can process in a single request. Frontier models now accept about 1 million tokens. Larger context windows allow more information but increase cost and latency, and models use very long inputs less reliably than short ones.See glossary: The model can consider a fixed amount of text at once. Leading models now accept up to about 1 million tokens per request -- thousands of pages -- though some smaller models still stop at around 200,000. Models use very long inputs less reliably than short ones, so a big window does not replace good retrieval
Reasoning ("Thinking") Models
The newest models add a step. Instead of answering in a single pass, a reasoning modelReasoning ModelA model that works through a problem step by step and checks itself before answering (also called a thinking model; the extra work is known as test-time compute). More accurate on analysis and multi-step questions, but slower and costlier because its thinking is billed as output.See glossary works through a problem step by step and checks itself before it responds -- the industry calls this "test-time compute." The result is better accuracy on analysis and multi-step questions, at the price of slower, costlier answers (the thinking is billed as output). Reasoning is now the default on flagship models. Think of an analyst who works the numbers on a scratch pad and checks them before signing the memo.
BANKING ANALOGY
Think of an LLM like your most experienced credit analyst -- the one who has reviewed thousands of loan applications over a 30-year career. When a new application arrives, they do not consult a rulebook for every decision. Instead, they draw on deep pattern recognition built from years of experience. They can spot risk factors, identify inconsistencies, and draft recommendation memos that sound authoritative. But critically, they are working from patterns in past data, not from real-time market feeds. An LLM operates the same way: powerful pattern recognition built from training data, but with no awareness of what happened after its training cutoff date.
What Makes LLMs Different from Traditional AI
Traditional AI in banking -- the kind powering your credit scoring models and fraud detection systems -- is narrow and task-specific. A fraud detection model does one thing exceptionally well. An LLM, by contrast, is a generalist. The same model that can summarize a 100-page regulatory filing can also draft a customer communication, explain a complex derivatives structure, or generate Python code for a risk calculation.
This generality is both the opportunity and the challenge. LLMs are remarkably versatile, but that versatility means they require careful governance -- a topic we will explore in depth in the Governance and Risk module.
Explore the interactive AI technology stack below to see where LLMs fit -- click any layer to see what it does, which tools power it, and how banks use it.
Swipe sideways to see the whole diagram
Click or use arrow keys to explore
Why Banking Executives Should Care
The banking industry generates and consumes more text than almost any other sector. Consider the daily volume of:
- Regulatory documents: Basel frameworks, OCC bulletins, Fed guidance, state-level regulations
- Customer communications: Emails, chat transcripts, complaint letters, advisory notes
- Internal reports: Credit memos, risk assessments, audit findings, board presentations
- Legal documents: Loan agreements, compliance certifications, litigation filings
- Market research: Analyst reports, earnings calls, economic forecasts
Every one of these text-heavy workflows is a potential LLM use case. Banks that deploy LLMs strategically can process information faster, reduce manual effort on routine tasks, and free their most expensive resource -- experienced professionals -- to focus on judgment-intensive work.
Key Capabilities for Financial Services
LLMs bring several capabilities that map directly to banking needs:
Document Summarization and Analysis
An LLM can read a 200-page regulatory proposal and produce a structured summary highlighting the provisions most relevant to your institution. What currently takes a compliance analyst two days can be reduced to minutes -- with the analyst then reviewing and refining the output rather than starting from scratch.
Intelligent Search and Retrieval
When combined with your internal document repositories (a technique called Retrieval-Augmented Generation, or RAG), LLMs can answer natural-language questions about your own policies, procedures, and historical decisions. Instead of searching through SharePoint folders, a relationship manager could ask: "What is our current policy on commercial real estate concentration limits?"
Reading Documents and Images
Current models read scanned pages, charts, and photos, not just typed text -- putting scanned loan packages, statement PDFs, and cheque images within reach. Extracted figures still need checking before they feed a credit decision.
Draft Generation
From customer correspondence to board reports to regulatory responses, LLMs can generate first drafts that human experts then refine. This shifts the workflow from creation to curation -- a significantly more efficient model.
Code and Data Analysis
LLMs can write SQL queries, Python scripts, and data analysis code from plain-English descriptions. This democratizes data access for business users who understand the questions but lack programming skills.
Limitations and Risks
No discussion of LLMs is complete without an honest assessment of their limitations -- particularly in a regulated industry like banking.
HallucinationsHallucinationWhen an AI model generates plausible-sounding but factually incorrect information. A critical risk in banking where inaccurate outputs could lead to regulatory violations or financial losses.See glossary. LLMs can generate plausible-sounding but factually incorrect information. In banking, where accuracy is non-negotiable, every LLM output touching customers, regulators, or financial decisions must be verified by qualified humans.
Data privacy. Sending sensitive customer data or proprietary trading strategies to a cloud-hosted LLM creates data leakage risk. Many banks deploy LLMs within their own infrastructure or use enterprise agreements with strict data handling provisions.
Bias and fairness. LLMs inherit biases present in their training data. If used in any capacity that influences lending, pricing, or customer treatment decisions, banks must validate for fair lending compliance -- with the same discipline they apply to traditional credit models under model risk management.
Regulatory uncertainty. In April 2026 the Federal Reserve, OCC and FDIC replaced SR 11-7 (and OCC Bulletin 2011-12) with updated model risk managementModel Risk ManagementSupervisory guidance and bank practices governing how models are inventoried, validated, monitored, and controlled. Since April 2026 the interagency guidance is SR 26-2 (with OCC Bulletin 2026-13 and FDIC FIL-15-2026), which replaced SR 11-7 and OCC Bulletin 2011-12. Generative and agentic AI are explicitly outside its scope, but examiners can still act on unsafe or unsound practices or violations of law, and most banks continue to apply model-risk disciplines to their AI systems.See glossary guidance -- SR 26-2, OCC Bulletin 2026-13 and FDIC FIL-15-2026. The new guidance is aimed mainly at banks with over $30 billion in assets, and it explicitly places generative and agentic AI outside its scope while the agencies gather input on how banks use AI (the OCC said a request for information is coming). That is not a free pass: examiners can still act on unsafe or unsound practices or violations of law, and most banks continue to apply model-risk disciplines -- inventory, validation, monitoring, documentation -- to their AI systems. Deploying LLMs without a clear governance framework is a material risk.
Tip
When evaluating LLM vendors for your institution, apply the same rigor you would to any critical technology vendor. Ask about data residency, model versioning, audit trails, and incident response. Insist on enterprise agreements that address financial services compliance requirements -- consumer data protection, model risk management, and third-party risk management. The cheapest API plan is rarely appropriate for banking workloads. And expect prices to keep moving: the cost of a given level of AI capability has been falling by roughly 9x to 900x a year depending on the task (about 50x a year at the median), so budget for capability rather than today's price list, and plan to re-price contracts often.
What Comes Next
Understanding what an LLM is represents the first step. In the units that follow, you will learn how different foundation models compare, how to connect LLMs to your own data through vector databases and RAG, and how to build the governance framework that makes responsible deployment possible.
The executives who understand these technologies well enough to make informed strategic decisions -- not to build them -- will guide their institutions through this transformation.
KNOWLEDGE CHECK
What is the most accurate way to describe how a Large Language Model generates responses?
A bank is evaluating where to deploy an LLM. Which use case introduces the MOST significant risk that requires human oversight?
Why is a large context window (now up to about 1 million tokens on frontier models) strategically important for banking applications?