A deal stalls on a word. Someone says token savings or vector database, the room nods, and half of it has quietly stopped following. The nodding is the trap: almost nobody announces that a term lost them, so the conversation runs on over a gap neither side can see. Of everything buyers ask us, the plainest request is the one they make most often: what does that actually mean?
AI Glossary: What Do
These Terms Actually Mean?
Plain-language definitions of the words that stop conversations — token, RAG, vector database, NPU, air gapped, hallucination — each with the enterprise consequence that makes the word worth knowing.
Asking what a term means is the question buyers raise more than any other. Across our sales and customer conversations, buyers interrupt to ask for a plain-language definition more often than they raise any other buyer-side subject. Several internal sales-operations subjects surface more often across the record as a whole, so this is the leading buyer question rather than the leading question overall — and none of it is search volume.
The limit: nobody can rank which individual terms confuse people most. Buyers phrase the definition request almost as many ways as they make it, and mined questions rarely recur from one room to the next, so term-level demand does not exist in our evidence. Two weaker signals do. Buyers have named terms they could not follow, and we can tally how often each term was spoken in our rooms — a tally that measures use rather than confusion, carries no audience label, and sets the reading order below without ranking difficulty.
Every entry runs two sentences: what the word means, then what it costs or buys you. Read the first to repeat the term in your own meeting. Read the second to see why the word is on the table at all, because most of these terms hide a decision — how fast is fast enough, how much can the model hold, what leaves the machine — and each entry points at the page that settles it rather than settling it here.
Vocabulary and capability are separate problems, and definitions only fix the first. A fluent speaker of AI terminology can still be helpless with the tool, which is why definitions sit apart from skills, curriculum and rollout design. For more information on the wider barriers to getting a workforce using AI, visit the workforce adoption pillar.
What Does LLM Stand For? The Acronyms, Expanded
LLM stands for large language model and SLM for small language model. The letters describe size and where the model runs: an LLM carries tens to hundreds of billions of parameters and needs a GPU server or an API, while an SLM of 100 million to 10 billion parameters fits on a laptop.
Some of the confusion is simpler than a definition. The letters get said out loud and never expanded, so the sentence moves on while half the room is still decoding three characters. Each entry below gives the expansion, one example and the page that goes deeper.
- LLM — large language model
- A general-purpose model trained on broad text to answer, summarize and draft in plain language — ChatGPT, Claude and Gemini all run on one. Tens to hundreds of billions of parameters means it is served from a GPU server or an API rather than from the machine in front of you. For more information visit the SLM and LLM comparison page.
- SLM — small language model
- The same technology at 100 million to 10 billion parameters, small enough to run on a laptop, a phone or an edge device; Microsoft Phi, Google Gemma, Qwen and Llama at 8B are named families. A smaller model answers locally and cheaply, and it holds less general world knowledge to answer from. For more information visit the SLM and LLM comparison page.
- Token
- The unit a model reads and writes, closer to a fragment of a word than a whole one, so a long word can arrive as two or three of them. Everything a model does is counted in tokens, which makes tokens what a hosted service bills for and what sets how quickly text appears locally. For more information visit the token and inference cost page.
- Context window
- How much text a model can hold in mind at once, counted in tokens and shared between your question, the material retrieved for it and the answer coming back. A long document can exceed the window, and the largest windows are out of reach on a laptop. For more information visit the sizing page.
- RAG — retrieval-augmented generation
- The model searches your own documents first, then writes its answer out of the passages it found rather than out of training. AirgapAI runs retrieval-augmented generation over the data set you select and shows the source citations behind each answer. For more information visit the retrieval architecture page.
- Temperature
- The setting that decides how much variation a model allows when it picks the next word. Keep it low when a policy answer has to come out the same way twice; raise it when you want several different drafts to choose between. For more information visit the accuracy page, or the guide to temperature, top-p and other sampling settings.
- Parameters
- The learned weights inside a model, quoted in billions, so an 8B model holds eight billion of them. Parameter count is the first thing that sets memory, because the larger the number, the more system or video memory the machine needs before the model will load at all. For more information visit the sizing page.
Token, retrieval-augmented generation and the context window each earn a fuller entry with the enterprise consequence attached, starting with the six words buyers said they could not follow.
Why Buyers Keep Stopping to Ask
The requests rarely sound technical. Buyers repeatedly stop us to ask whether an AI hallucination means the tool spits out wrong information, whether FTE means full-time employee, whether audience segments means demographics such as age and income, and whether router refers to the one inside the local app doing the reasoning or the one routing a query out to a server. Partners ask whether activation means activating with the sales force, the product side or operations.
Two patterns run through all of them. The first is ordinary English pressed into specialist duty: router, activation, client, tenant, agent, block. Each word already means something to a business audience, and the technical meaning arrives without warning. The second is the compound acronym — RAG, NPU, DSR — which carries a social cost to admit missing, so people carry on and hope context rescues them. Both patterns are translation problems, and translation is the job of the list below.
Which Terms Earn an Entry, and What the Evidence Cannot Rank
Two sources decide what earns an entry, and both have limits worth stating before you read a single definition.
- Terms buyers named as unfollowable. A recurring complaint names specific words that lost an audience. Those named words are the strongest evidence available for what to define first, so they lead the glossary.
- How often a term was spoken. We can tally the terms of art used across our recorded conversations and order the rest of the glossary by that tally. It measures how often a word was used, never how often it was asked about, and it does not separate who said it — so it is neither a list of what buyers find confusing nor search volume.
Neither source supports a ranking. The definition request arrives in a wide spread of one-off phrasings and mined questions rarely repeat across rooms, so the evidence does not support a league table of the most-misunderstood terms. Any glossary claiming one is guessing.
The selection rule follows. In the two ordered lists below, a term earns its place only when the record carries it, so words that felt obvious but never showed up stayed out of them. Product nouns are spelled the way people actually say them, which is why you will read idea blocks below rather than a tidier engineering spelling.
What the Confusion Sounds Like in the Room
One exchange shows the shape of the problem better than any taxonomy. A buyer heard the word block inside Blockify and asked whether the product uses block data or block storage, the kind that cannot be tampered with or hacked. The answer is no. Blockify output is text: idea blocks hold structured text pulled out of your own source documents, and security comes from tagging and metadata governance rather than from any storage format.
Treat that as an illustration rather than a common misconception; it came from one buyer, in one room. What it shows is how the collision happens: an English word gets welded onto a product name, the buyer reads the familiar meaning, and both sides leave the call believing they agreed. The same trap waits inside every product noun built from a word people already own.
The Vocabulary Gap Runs Both Ways
Buyers describe the gap in both directions, and the second direction is the uncomfortable one. Board members know all the AI terminology without knowing what any of it means. A technical leader talks about vector databases and the customer glazes over. Experts reach for tokens, vector databases and JSON in front of people who follow none of the three, and buyers who had never had a token explained heard a token-savings claim as noise. Most people, buyers tell us, still do not know what RAG is or what an AI agent is. Even the categories collide: models run on both the AI factory and the AI data platform, so the boundary between them reads as unclear.
Those complaints name six words outright — token, vector database, RAG, AI agent, JSON, and the AI factory and AI data platform pair. Named terms beat inferred ones, so those six lead the glossary; everything after them is ordered by how often the term was spoken.
The Six Words Buyers Said They Could Not Follow
Each entry gives the meaning in one sentence you can repeat, then the enterprise consequence in a second, then the page that goes deeper.
- Token
- A token is the unit a language model reads and writes, roughly a fragment of a word, and every model measures and prices its work in tokens. Token counts drive the bill on a hosted service and the response speed you feel on a local machine, which is why a buyer who cannot picture a token cannot judge a token-savings claim. For more information visit the token and inference cost page.
- Vector database
- A vector database stores passages of text as numeric coordinates so a search can retrieve them by meaning instead of by exact wording. Retrieval quality is won or lost there, and Iternal states that adopting Blockify requires no change to an existing vector database, re-ranker or document parser, because Blockify sits between the chunks and the store. For more information visit the retrieval architecture page.
- RAG, or retrieval-augmented generation
- Retrieval-augmented generation searches your own documents first, then hands the passages it found to the model to compose the answer. RAG is how a local assistant answers questions about your contracts without the model ever having been trained on them: AirgapAI runs retrieval-augmented generation over the selected data set and shows the source citations behind each answer. For more information visit the retrieval architecture page.
- Foundation model
- A foundation model is the general-purpose model trained on broad public material before anyone points it at your work, the base layer a retrieval system or a fine-tune builds on. Which one you run is a procurement decision rather than a fixed property of the software: AirgapAI loads the model you choose and sizes it to the hardware in front of you. For more information visit the bring your own model page.
- AI agent
- An agent is one specific action an AI takes on your behalf, such as searching a data set or working a saved sequence of steps, and in practice it is often little more than a prompt held in a text file. Agents become a governance question the moment they act, because an agent gets tool access provisioned the same way access is provisioned for a person. For more information visit the agents and skills page.
- JSON
- JSON is a plain-text format that stores structured data as labeled fields, legible to a person and a program alike. JSON is how prepared data and AI configuration travel between tools: Iternal ships blockified data sets as lightweight JSONL packages, and AirgapAI workflows are JSON files you import. For more information visit the components page.
- AI factory and AI data platform
- An AI factory is the concentrated compute estate an organization stands up to run models at scale, carrying cost across data center, services, maintenance, uptime and network security; an AI data platform is the storage and orchestration layer above it that governs which content those models may read. Models run on both, which is why the boundary blurs, so use the blunt test: one holds your GPUs, the other holds your documents. For more information visit the placement page.
The Terms Spoken Most Often in Our Rooms
The order below follows how often each term was used in conversation, which tracks what gets talked about rather than what puzzles anyone. Read it as a reading order, never as a difficulty ranking.
- Air gapped
- An air-gapped machine has no path to the public internet at any stage of its life: software, model, data and every later update reach it by hand, on removable media or from a server inside your own walls. The phrase gets used loosely, so test it rather than accept it, because running the model on the device is a weaker standard than never having a route out. For more information visit the offline and air-gapped page.
- Sovereign AI
- Sovereign AI means non-public, non-hyperscale AI: models and corpora an organization or a country owns and operates inside its own boundary. Iternal designed AirgapAI as local sovereign AI a company controls within its own walls. Sovereignty is a question of who can switch you off, because one provider withdrawing access removes an entire agent workforce at once. For more information visit the data residency page.
- Idea blocks
- Blockify carves PDFs, Word files and other source material into idea blocks, each holding one critical question and its trusted answer, tagged with metadata describing what it answers and which departments it serves. A block never mixes two ideas, which is where the accuracy comes from, and the metadata is what makes an answer permissionable and traceable to its source. For more information visit the retrieval architecture page.
- Tool calling
- Tool calling is a model reaching outside itself to run an action, read a record or write to another system. Iternal is direct about where the local edge of that capability sits today: AirgapAI delivers day-to-day productivity without tool calling, and smaller local models remain less reliable at it than larger server-class models. For more information visit the integration page.
- NPU
- A neural processing unit is a low-power chip built into newer laptops specifically for AI work, sitting alongside the CPU and the GPU. Iternal targets a laptop with an NPU as the right machine for AirgapAI and states the current limit plainly: NPU models load slowly from cold, a delay that lives inside the chipmaker inference engine rather than in the application. For more information visit the sizing page.
- Chunking
- Chunking splits a document into fixed-size pieces so a retrieval system can index them. Arbitrary chunking cuts text off at a character count and hands the model fragments stripped of their context, the failure Blockify exists to remove; AirgapAI still offers basic chunking as a fast alternative, and it holds up when one person controls ten or fifteen files with no version sprawl. For more information visit the retrieval architecture page.
- Hallucination
- A hallucination is an AI answer that is fluent, confident and wrong — in a buyer phrasing we hear often, the tool spits out wrong information. Iternal traces most enterprise hallucination to messy source data and states that Blockify virtually eliminates it: virtually, because the share caused by the model intelligence itself sits outside any product to fix. For more information visit the accuracy page.
- Context window
- The context window is how much text a model can hold in mind at once, counted in tokens. The window caps what you can attach, since a hundred-page PDF can exceed it, output gets less accurate as a long window fills, and running locally puts the largest million-token windows out of reach; AirgapAI benchmarks the hardware it finds and recommends a context window to match. For more information visit the sizing page.
- Tokens per second
- Tokens per second is the rate at which a model produces text on the hardware in front of you. Iternal treats roughly thirty tokens per second as usable throughput and fifteen as not, with output at thirty-five to forty faster than most people read; run the model locally and the figure gauges response speed rather than cost. For more information visit the benchmarks page.
The Words That Stretch, and How to Pin Them Down
Several words above describe a range rather than a point. Air gapped, sovereign, agent and platform each cover a spread of real architectures, so two suppliers can each use one of them accurately and still mean different things by it. The fix is procedural: make the definition written, specific and attached to the paperwork. Iternal answers each of the following in writing, and each is worth putting to anyone quoting you a private AI system.
-
Define air gapped for the deployment you are quoting: which lifecycle steps assume a network connection?Whether you are buying a machine that tolerates losing the network or one that never had a route out.
-
Define sovereign for this agreement: which jurisdiction holds the compute, the corpus and the keys?Whether sovereignty is a property of the architecture or a description of a hosting address.
-
When you say agent, do you mean a saved prompt, a retrieval action, or a loop that reads and writes to our systems?The scope of the access review, which changes completely across the three.
-
Attach a one-sentence definition of every term of art in the proposal to the statement of work.That both sides mean the same thing by the same word before any money moves.
One discipline keeps a glossary useful. A definition names the page that goes deeper and stops, because a glossary that tries to teach retrieval architecture inside a definition has stopped being a glossary. Hold the same discipline in your own documents: one sentence of meaning, one sentence of consequence, one link.
- Turning the right words into a useful answer from the model — see the prompting page.
- What a role-by-role AI training plan should teach, and to whom — see the training curriculum page.
- Moving an organization from a pilot to a daily habit, including after a rollout that went badly — see the adoption behavior page.
- How the ingestion engine, the data set and the local assistant connect end to end — see the components page.
- What your people may and may not type into an AI tool — see the acceptable-use page.
FAQ: AI and Deployment Terminology
LLM stands for large language model: a general-purpose AI model trained on a broad body of text to answer questions, summarize documents and draft language. The systems behind ChatGPT, Claude and Gemini are all large language models. Size is the part the name points at, tens to hundreds of billions of parameters, which is why an LLM is served from a GPU server or an API rather than from the laptop in front of you.
SLM stands for small language model: the same technology built at 100 million to 10 billion parameters so it can run on a laptop, a phone or an edge device. Microsoft Phi, Google Gemma, Qwen and Llama at 8B are named families. The trade is capability for locality. A small model answers on the machine in front of you, at low cost and without a network, and it holds less general world knowledge than a frontier model.
It means the model produced an answer that reads fluently, sounds certain and is wrong. Buyers put it more bluntly when they ask whether a hallucination means the tool spits out wrong information, and the answer is yes. Iternal locates most enterprise hallucination in the source data rather than the model, which is why Blockify pre-processes documents and why Iternal says it virtually eliminates hallucination: the share caused by model intelligence sits outside any product to fix.
A neural processing unit is a third chip in newer laptops, alongside the CPU and the GPU, designed to run AI work at low power. Iternal treats a laptop with an NPU as the right target machine for AirgapAI. The trade-off today is load time: NPU models start cold slowly, and that delay sits inside the chipmaker inference engine rather than in the application, so plan around it.
A token is the unit a model reads and writes, roughly a fragment of a word. Everything a model does is measured in tokens, so tokens are what a hosted service bills you for and what sets how quickly text appears on a local machine. Buyers who have never had a token explained hear a savings claim as noise, which is what happened while the pricing was subsidized and nobody had to care.
RAG is the method; the vector database is one of the parts it uses. Retrieval-augmented generation searches your own material first, then hands what it found to the model to compose an answer, and the vector database is the store holding your text as numeric coordinates so that search can work by meaning instead of exact wording. AirgapAI runs retrieval-augmented generation over the data set you selected and shows the source citations behind each answer.
No. A buyer asked us exactly that after hearing the word block, wondering whether the data was held in something tamper-proof, and the answer is that Blockify output is ordinary text. Idea blocks are structured text lifted out of your source documents, each pairing a critical question with its trusted answer; the protection comes from tagging and metadata governance, not from the storage layer.
Because the vocabulary gap runs in two directions. Board members can hold all the terminology without knowing what any of it means, while specialists reach for tokens, vector databases and JSON in front of audiences who follow none of them. Even the categories collide, since models run on both an AI factory and an AI data platform and the boundary reads as unclear. The cure is a written definition per term, agreed once, in the paperwork.
Put the Definitions in the Paperwork
Vocabulary is cheap to fix and expensive to leave alone. Take the dozen terms your own project turns on, write one sentence of meaning and one of consequence for each, and attach the list to the statement of work. Words stop drifting once both sides have signed the same definitions, and the meeting that used to stall on what does that mean gets its hour back.