Prompts, Completions, and System Instructions
Jump to a section
The Interface Layer: How You Talk to an LLM
In traditional banking software, you interact through structured forms, dropdown menus, and predefined workflows. A loan origination system does not ask for your opinion -- it asks for a borrower's income, credit score, and collateral value in precisely defined fields.
Large Language ModelsLarge Language Model (LLM)A neural network trained on vast amounts of text data that can understand and generate human language. LLMs power chatbots, document analysis, code generation, and many enterprise AI applications.See glossary work differently. Instead of structured inputs, you communicate in natural language -- plain English sentences and paragraphs. This interface layer, built on prompts, completions, and system instructions, is what makes LLMs both remarkably accessible and surprisingly nuanced to use well.
Understanding this interface is essential for any banking executive evaluating AI tools, because the quality of your instructions directly determines the quality of the output. A poorly prompted LLM is like a brilliant but poorly briefed analyst: capable of excellent work, but likely to deliver something you did not actually need.
KEY TERM
Prompt: The text input you provide to an LLM -- your question, instruction, or request. A prompt can be a single sentence ("Summarize this regulatory filing") or a detailed multi-paragraph instruction set with examples, constraints, and formatting requirements.
Prompts: Your Instructions to the Model
A promptPrompt EngineeringThe practice of crafting effective instructions (prompts) to guide AI model behavior. Techniques include few-shot examples, chain-of-thought reasoning, and role-based system instructions.See glossary is simply the text you send to an LLM. When a relationship manager types "Draft a follow-up email to a commercial lending prospect who toured our treasury management platform," that entire sentence is the prompt. The model reads it, interprets the intent, and generates a response.
But not all prompts are created equal. The emerging discipline of prompt engineering focuses on crafting instructions that consistently produce high-quality, reliable outputs. For banking applications where accuracy and tone matter, this discipline is critical.
Zero-Shot vs. Few-Shot Prompting
Zero-shot prompting means giving the model a task with no examples. You simply describe what you want:
"Classify this customer complaint as one of: billing dispute, fraud report, service quality, or account access."
Few-shot prompting means including examples in your prompt to demonstrate the pattern you expect:
"Classify customer complaints. Examples:
- 'I was charged twice for my wire transfer' -> billing dispute
- 'Someone opened a card in my name' -> fraud report Now classify: 'Your mobile app has been down for three days'"
Few-shot prompting typically produces more consistent results for banking tasks because it removes ambiguity about your expectations. When classifying regulatory correspondence or tagging transaction categories, a few well-chosen examples dramatically improve accuracy.
BANKING ANALOGY
Think of prompting like briefing a new analyst on your team. If you say "Review this loan file," you might get anything from a one-paragraph summary to a 20-page analysis. But if you say "Review this loan file, focusing on the three largest risk factors, and present your findings in a one-page memo formatted like the example I am attaching," you will get exactly what you need. The same principle applies to LLMs -- specificity and examples produce better results.
Completions: What the Model Returns
A completion is the model's response to your prompt. The term comes from the original framing of LLMs as text-completion engines: given a sequence of tokensTokensThe basic units of text that LLMs process. A token is typically one-half to three-quarters of an English word, depending on the model. Token counts determine cost, speed, and context window limits for every API call.See glossary, the model "completes" the sequence by predicting what comes next.
In practice, completions can be anything from a single word to a multi-page document, depending on your prompt and configuration. For banking applications, completions might include draft regulatory responses, summarized credit memos, classified customer inquiries, or generated code for data analysis.
System Instructions: Setting the Ground Rules
System instructionsSystem InstructionsPersistent instructions provided to an LLM that define its role, behavior constraints, and output format. System instructions shape every response without being visible to end users.See glossary are persistent directives that shape every response the model generates. Unlike a prompt -- which changes with each query -- system instructions remain constant across an entire session or application. They define the model's role, behavioral constraints, and output format.
KEY TERM
System Instructions: A special category of prompt that sets persistent behavioral guidelines for the LLM. System instructions are processed before every user prompt and define the model's persona, constraints, tone, and output requirements. End users typically do not see system instructions.
For banking applications, system instructions are where you encode compliance requirements, tone guidelines, and safety constraints. For example:
- "You are a compliance assistant for a US-based commercial bank. Never provide legal advice. Always cite the specific regulation when referencing regulatory requirements. If uncertain about any regulatory interpretation, state that explicitly."
This single system instruction transforms a general-purpose LLM into a focused, appropriately constrained banking tool. Treat it as a control, not a guarantee: users -- or documents the model reads -- can try to override system instructions, an attack called prompt injectionPrompt InjectionAn attack in which instructions hidden in user input, or in content an AI reads (an email, document, or web page), try to override its rules. Ranked the top risk in the OWASP 2026 Top 10 for LLM applications. The main defenses are least-privilege permissions and human approval for sensitive actions.See glossary that tops the OWASP 2026 list of LLM risks. The Agents unit covers the guardrails that back it up.
BANKING ANALOGY
System instructions are like the compliance guidelines you give a new analyst on their first day. Before they answer a single client question, they learn: "We never guarantee returns. We always disclose fees. We refer legal questions to General Counsel." These ground rules shape every interaction without needing to be repeated each time. System instructions work the same way for an LLM.
Controlling Output: Settings, Formats, and Token Limits
A few settings and techniques give you control over how the model generates responses -- and they have changed as models have.
Reasoning Effort and Temperature
Older models expose a temperatureTemperatureA parameter controlling the randomness of model outputs. Lower temperature (0.0-0.3) produces more focused, repeatable responses; higher temperature (0.7-1.0) produces more creative, varied outputs. Many newer reasoning models do not expose temperature and use a reasoning-effort setting instead, and even temperature 0 does not guarantee identical outputs.See glossary dial that controls the randomness of the output, usually from 0 to 1:
- Low temperature (0.0 - 0.3): More focused, repeatable responses -- the traditional choice for classification, data extraction, and compliance checks
- High temperature (0.7 - 1.0): More varied, creative responses -- useful for brainstorming or generating diverse draft options
Many newer reasoning modelsReasoning ModelA model that works through a problem step by step and checks itself before answering (also called a thinking model; the extra work is known as test-time compute). More accurate on analysis and multi-step questions, but slower and costlier because its thinking is billed as output.See glossary do not expose temperature at all. Instead they offer a "reasoning effort" setting that controls how much the model thinks before answering; on Anthropic's current models, for example, a non-default temperature is rejected outright. And even at temperature 0, cloud AI services do not guarantee identical output for identical input, because of how their servers batch requests together.
The practical lesson: you cannot buy consistency with a single dial. For banking tasks where the same input should produce the same answer, consistency comes from clear system instructions, few-shot examples, a required output format, a pinned model version, and ongoing testing.
Structured Outputs
Models can now be required to answer in a fixed, machine-readable format -- for example, named fields for borrower, loan amount, and risk rating. With structured outputsStructured OutputsA feature that forces a model to answer in a fixed, machine-readable format, such as named fields for borrower, amount, and risk rating. It guarantees the format, not the accuracy of the values.See glossary, extraction from loan documents feeds downstream systems reliably instead of arriving as free text someone has to re-key. One caution: the format is guaranteed; the facts inside it are not. Validate the values the same way you would check a junior analyst's spreadsheet.
Token Limits and Cost
Every LLM has a context windowContext WindowThe maximum amount of text (measured in tokens) a model can process in a single request. Frontier models now accept about 1 million tokens. Larger context windows allow more information but increase cost and latency, and models use very long inputs less reliably than short ones.See glossary -- the maximum number of tokens it can process in a single request, covering both your prompt and the model's response. With frontier models now accepting about 1 million tokens, a 200-page regulatory filing fits easily.
The practical constraint today is cost. Every input token is billed on every request, so sending the same long document with each question adds up quickly; reuse (caching) and retrieval of only the relevant passages still matter. Reasoning models add another line item: their thinking tokens count toward output limits and are billed as output.
Tip
When deploying LLMs for banking workflows, establish standard system instructions for each use case -- compliance review, customer communication drafting, document summarization. Version-control these instructions the same way you version your compliance policies. This ensures consistency across your organization and creates an audit trail for regulators. Pin the model version too: vendors retire models on published schedules, and a model change can shift outputs as much as a prompt change -- treat it as change management.
Quick Recap
- Prompts are your natural-language instructions to the LLM -- their quality directly determines output quality
- Completions are the model's generated responses, produced one token at a time
- System instructions set persistent behavioral guidelines that shape every response, functioning like compliance rules for the AI
- Few-shot prompting (providing examples) produces more consistent results than zero-shot for banking classification tasks
- Consistency comes from instructions, examples, required output formats, pinned model versions, and testing -- not from a single setting; many reasoning models replace temperature with a reasoning-effort control
- Structured outputs guarantee the format of an answer, not its accuracy, and token usage drives cost even now that long documents fit in one request
KNOWLEDGE CHECK
A bank wants its compliance tool to classify regulatory correspondence consistently. Which approach is most appropriate?
What is the primary purpose of system instructions in a banking AI application?
Why is few-shot prompting particularly valuable for banking use cases compared to zero-shot prompting?