Glossary
A
- Agent2Agent (A2A)
- An open standard that lets AI agents from different vendors communicate and hand work to each other. It reached version 1.0 in early 2026 and joined the Linux Foundation's Agentic AI Foundation in August 2026. It complements MCP, which connects agents to tools and data.
- Agents
- AI systems that can autonomously plan and execute multi-step tasks by calling tools, querying data sources, and making decisions without human intervention at each step -- typically within defined permissions and human approval checkpoints.
- API (Application Programming Interface)
- A standardized interface that allows software systems to communicate. In AI, APIs let your applications send prompts to a model and receive generated responses programmatically.
C
- Chunking
- The process of splitting large documents into smaller, overlapping segments before generating embeddings. Chunk size and overlap strategy directly affect retrieval quality in RAG systems.
- Computer-Use Agent
- An AI agent that operates software the way a person does -- reading the screen, clicking, and typing. Useful for legacy systems with no modern interface, but it needs the strictest guardrails because on-screen content can carry injected instructions.
- Context Window
- The maximum amount of text (measured in tokens) a model can process in a single request. Frontier models now accept about 1 million tokens. Larger context windows allow more information but increase cost and latency, and models use very long inputs less reliably than short ones.
E
- Embeddings
- Numerical representations (vectors) of text that capture semantic meaning. Similar concepts produce vectors that are close together, enabling machines to understand relationships between words, sentences, or documents.
F
- Fine-Tuning
- The process of further training a pre-trained model on a specific dataset to specialize its behavior for a particular domain or task, such as banking compliance language.
- Foundation Model
- A large AI model trained on broad data that can be adapted to many tasks. Examples include OpenAI's GPT, Anthropic's Claude, Google's Gemini and Meta's Llama families. Banks evaluate these for capabilities, safety, and regulatory fit.
G
- Guardrails
- Safety mechanisms that constrain AI model outputs, and the actions AI agents can take, to prevent harmful, off-topic, non-compliant, or unauthorized results. Critical in banking for regulatory adherence and brand safety.
H
- Hallucination
- When an AI model generates plausible-sounding but factually incorrect information. A critical risk in banking where inaccurate outputs could lead to regulatory violations or financial losses.
- Hybrid Search
- Search that combines keyword matching with meaning-based (vector) search, often followed by re-ranking, where a second model re-scores the top results. It catches exact identifiers such as account numbers and regulatory citations that pure vector search can miss, and is current best practice for RAG.
I
- Inference
- The process of running a trained model to generate predictions or outputs from new input data. Inference cost, latency, and throughput are key factors in enterprise AI deployment.
L
- Large Language Model (LLM)
- A neural network trained on vast amounts of text data that can understand and generate human language. LLMs power chatbots, document analysis, code generation, and many enterprise AI applications.
M
- Model Context Protocol (MCP)
- An open standard for connecting AI models and agents to tools and data sources, so a connector built once works with any compatible model. Governed by the Linux Foundation's Agentic AI Foundation since December 2025. Each MCP connector is a new access path into bank systems and needs access and third-party risk review.
- Model Risk Management
- Supervisory guidance and bank practices governing how models are inventoried, validated, monitored, and controlled. Since April 2026 the interagency guidance is SR 26-2 (with OCC Bulletin 2026-13 and FDIC FIL-15-2026), which replaced SR 11-7 and OCC Bulletin 2011-12. Generative and agentic AI are explicitly outside its scope, but examiners can still act on unsafe or unsound practices or violations of law, and most banks continue to apply model-risk disciplines to their AI systems.
- Multimodal
- Describes models or embeddings that handle images, scanned documents, and audio as well as text -- for example, reading a scanned loan package or a cheque image.
O
- Orchestration Framework
- Software that coordinates LLMs, tools, and data sources into complex workflows. Frameworks like LangGraph, CrewAI, and vendor toolkits such as the OpenAI Agents SDK and Microsoft Agent Framework manage prompt chains, memory, and tool calling for multi-step AI tasks.
P
- Prompt Engineering
- The practice of crafting effective instructions (prompts) to guide AI model behavior. Techniques include few-shot examples, chain-of-thought reasoning, and role-based system instructions.
- Prompt Injection
- An attack in which instructions hidden in user input, or in content an AI reads (an email, document, or web page), try to override its rules. Ranked the top risk in the OWASP 2026 Top 10 for LLM applications. The main defenses are least-privilege permissions and human approval for sensitive actions.
R
- Reasoning Model
- A model that works through a problem step by step and checks itself before answering (also called a thinking model; the extra work is known as test-time compute). More accurate on analysis and multi-step questions, but slower and costlier because its thinking is billed as output.
- Retrieval-Augmented Generation (RAG)
- A pattern that combines document retrieval with LLM generation. The system searches a knowledge base for relevant context, then feeds it to the model to produce grounded, accurate answers.
S
- Semantic Search
- Search that understands meaning rather than just matching keywords. Uses embeddings to find conceptually similar documents even when they use different terminology.
- Structured Outputs
- A feature that forces a model to answer in a fixed, machine-readable format, such as named fields for borrower, amount, and risk rating. It guarantees the format, not the accuracy of the values.
- System Instructions
- Persistent instructions provided to an LLM that define its role, behavior constraints, and output format. System instructions shape every response without being visible to end users.
T
- Temperature
- A parameter controlling the randomness of model outputs. Lower temperature (0.0-0.3) produces more focused, repeatable responses; higher temperature (0.7-1.0) produces more creative, varied outputs. Many newer reasoning models do not expose temperature and use a reasoning-effort setting instead, and even temperature 0 does not guarantee identical outputs.
- Tokens
- The basic units of text that LLMs process. A token is typically one-half to three-quarters of an English word, depending on the model. Token counts determine cost, speed, and context window limits for every API call.
- Transformer
- The neural network architecture underlying modern LLMs. Transformers use self-attention mechanisms to process relationships between all parts of the input simultaneously, enabling powerful language understanding.
V
- Vector Database
- A specialized database optimized for storing and querying high-dimensional vectors (embeddings). Enables fast similarity search across millions of documents for RAG and recommendation systems.
Z
- Zero-Shot / Few-Shot Learning
- The ability of LLMs to perform tasks with no examples (zero-shot) or just a few examples (few-shot) provided in the prompt, without requiring model retraining.