The machine votes first. A model that will not load on the laptops your people carry is not a candidate, whatever it scores on a public leaderboard, and that one constraint shortens the list faster than any criterion you will write. Executives ask which model is best. Their engineer asks something narrower: which of these will run here, and is it clever enough for the work?
Which LLM Models Are Supported,
and Can You Bring Your Own?
The catalog as it stands, a selection rule you can run on your own documents, and where every decision that follows the model choice gets settled.
Two questions, and the answer to the second one is yes. Iternal curates a library of open-weight models tested and validated to run well on local hardware, cataloged by the equipment being run and qualified alongside Intel, Dell and other OEMs; running any of them carries no additional fee. AirgapAI also supports bring-your-own-model by default, so a model your own security team cleared can be added under Settings instead. Pick the smallest model that clears your accuracy bar on your own documents, never the largest one the machine will technically load.
The limit: the smallest model that fits the machine is often not a model that clears the bar. Iternal says where the small end gives out rather than waiting to be asked. Output quality degrades on lower-tier hardware and with smaller models, and a small local model traverses a large folder of data slowly and answers at lower quality. Agentic tool calling has been the weak spot for small local models measured against server-class ones, and the tool-friendly small models are only now arriving for device-side AI. For code specifically, no model small enough for a current AI PC is genuinely useful. The general-purpose models do not cover every use case, so more specialized ones sometimes have to be layered in. And every shipped local model carries a training cut-off, so it may not know recent events.
Three things to establish before you standardize. Which models your build installs in one click and which download afterwards. Which models Iternal has validated on the class of machine your people carry. And what separates the catalog entries, per model. The targeted questions below turn each into a written answer.
Selection is a constrained trade, and the constraint is the device. The catalog and the selection procedure come first. Tuning, provider pricing, languages and speech, agents, building on the engine, systems integration and coding agents each have a page below that argues them out.
The Catalog as It Stands
Iternal maintains a curated library of open-weight models tested and validated to run well on today’s AI PCs, indexed by the equipment in front of the user. Gemma sits in it beside the Llama, Qwen and Mistral families; some entries are tuned for computational analysis, one for legal-sector work.
What ships inside the installer is small, and that is packaging rather than advice. A Windows installer is hard-capped at two gigabytes, so only a model around a billion parameters rides along. Iternal states the constraint plainly: the right model depends on the device, and defaulting everyone to the one-billion size would leave most people well below what their hardware could carry. Some packages ship Llama at one and three billion parameters; others hand the user a short written guide instead. Confirm what your build carries before you plan an offline rollout.
| Size class | What it is for | What it expects of the machine |
|---|---|---|
| About 1 billion | The only size small enough for a single-click install. Iternal runs one this size on device for data-set creation. | Almost anything, though 2023-era hardware is marginal. |
| About 3 billion | The Meta Llama Iternal advises customers to start testing with; strong on summarization and transcription. | 2024-generation AI PCs and newer; 16 GB of memory runs it, 32 GB is comfortable. |
| Four to nine billion | Where local deployments live and most customers settle: a Llama 3 in the 8-billion class, a Qwen 9B build for macOS. Below the catalog Llama, accuracy falls short of what Iternal expects. | A current AI PC with a capable integrated or discrete GPU. |
| 26 to 31 billion | Recent Gemma and Qwen releases small enough to run locally with reliable tool calling; one is a mixture of experts with about three billion active. | Workstation hardware, a desktop GPU, or a model served from your own data center. |
Two facts about that table sit together rather than apart. Most users never push a workload past the three-billion model, which is why the product leads with the sizes everyone can run; most deployments still settle on an eight-billion Llama. Need and practice are different measurements. Where the general-purpose entries miss a use case, Iternal packages a specialized model as a services engagement.
Set the Bar, Then Buy the Least Model That Clears It
Model choice goes wrong in a predictable way: somebody picks the biggest thing that loads, finds it slow, and blames local AI. The procedure below inverts that.
- Write the bar down before you look at a model. Name the task, then set the standard against the person who does it today rather than against an absolute. A draft an attorney will sign off is a different bar from a filing.
- Start at the smallest model your fleet can run. Every step up costs load time, memory and the number of machines you can deploy to. Model selection follows the use case, because no single model fits every job.
- Score it on your own documents. Published numbers will not settle this. Benchmarks are plentiful, unstandardized, and their scores do not translate into a meaningful capability level; a quantized result read against a full-precision baseline is not like-for-like either.
- Stop at the first model that clears the bar and keep the runner-up on the list, because the next hardware refresh may make it free.
Two habits keep the verdict from going stale. Re-run the test when a model generation turns over: Iternal describes model intelligence roughly doubling every six months while the size needed to reach a level halves. And plan to do the updating yourself — AirgapAI does not refresh local models on its own. An out-of-date model still works. Older is not broken.
Build the Test Set Nobody Handed You
The hardest step for a new user is not the install. It is working out which model to run. Buyers described the gap in almost the same words each time: comparing local models is hard without documentation of the differences between them, and a person installing manually has to work out which to load and what to expect from each. Iternal ships short descriptions in the app and longer write-ups on request. Neither replaces a test set built from your own work.
What a harness is, concretely. Twenty to forty real questions taken from work people already do, each with an answer you know is right, held in a file that outlives the evaluation. Run every candidate against it, score the answers, keep the file. When a release lands, you re-run rather than re-argue. One check the software will not run for you: AirgapAI still lets a user select a model the machine cannot carry, because the blocking flag is not in the build yet. Confirm the model fits before you read anything into its answers.
-
Which models does our build install in one click, and which download afterwards?Whether a machine with no route out is productive on day one.
-
Which catalog models has Iternal validated on the class of machine our people carry?The shortlist, narrowed by evidence rather than by parameter count.
-
Which models has Iternal validated for reliable tool calling, and on which hardware tier was that measured?Whether agentic work is on the table on the device, or belongs on a server you run.
-
Will Iternal put the differences between catalog entries in writing, model by model?The guesswork, replaced by a document your architects can argue with.
The Corpus Decides How Much Model You Need
Model size is the expensive lever. Corpus quality is the cheap one, and it moves the same number. Iternal argues the point about its own product line: cleaner input lets a smaller model do the same job. A well-structured set of documents supplied at answer time carries a smaller model over a bar an unstructured set would have needed a larger model to reach — which shows up as fewer machines to upgrade, not merely a better answer.
Read that as direction rather than discount. The figures underneath it were measured on specific corpora by a specific method and mislead when quoted alone, so they belong to the page that owns them. For more information visit the accuracy and traceable answers page.
Bringing In a Model You Chose Yourself
AirgapAI supports bring-your-own-model by default: any open-source model can be added under Settings, and a drop-down switches between the models you have installed. Iternal built it that way for a reason buyers supplied — approving a new model inside a large enterprise is very difficult, so the assistant takes the one your security team already cleared.
The build format is the real gate. Models arrive as llama.cpp or OpenVINO builds depending on the hardware, and a machine with no discrete GPU is offered the OpenVINO ones. OpenVINO is the Intel engine packaged inside AirgapAI: it runs across CPU, integrated GPU and NPU, optimizes to int4 rather than fp16 so a model takes far less video memory, and works only on Intel silicon. The generic engine will run models pulled from Hugging Face, Qwen, Gemma, Meta Llama and Mistral among them, and onboarding wants an embeddings model from the same folder as the language model. Models are folders of files in a known directory, so a management tool can swap one for another and the application picks it up on next load.
What is unsupported, said plainly. Converting and packaging a model yourself is a command-line process Iternal puts within reach of about one percent of customers. You may install as many models as you like, but one is active at a time. AirgapAI connects only to localhost by design, so a hosted model is not reachable from the chat application — the route to a larger model, or to tool calling, is a connection to a server you run, fronted by your own identity provider. Running locally also means no access to the largest models with million-token context windows.
One Model Choice, Seven Decisions Downstream
Picking the model settles less than it feels like. Every row below is a decision the choice sets up rather than closes, and each has its own page.
| What you are actually asking next | Where it is answered |
|---|---|
| You want the model to absorb your material permanently — and to know who owns the weights and whether its origin survives your security team. | the tuning and provenance page |
| You are exposed to a price somebody else sets and to capability shipped on somebody else’s schedule. | the provider economics page |
| Your people work in more than one language, or speak rather than type. | the languages and speech page |
| You want the assistant to act rather than answer — and to know what governs it when it does. | the agents, personas and skills page |
| The model is a component inside software you are building, your own interface on top. | the engine-behind-your-product page |
| The assistant has to read from the systems you already run and, eventually, write back. | the systems-integration page |
| Your engineers want an agent inside the repository, with guardrails and a spend ceiling. | the coding-agent practice page |
One path ends somewhere other than a page. When no model clears your bar on the machines you own, change the machines or change the bar — deliberately, rather than three months into a pilot.
- What a machine needs to carry a given model, and how to size a fleet — see the architecture and sizing page.
- How to measure throughput once a model is installed — see the speed measurement page.
- Whether the model belongs on the device, a server you own, or somewhere hosted — see the placement page.
- What a local deployment costs, and how the software is licensed — see the cost page.
- Getting the documents ready before any model reads them — see the data readiness overview.
FAQ: Picking a Model That Will Run on Your Machines
Iternal curates a library of open-weight models tested and validated to run well on local hardware, indexed by the equipment in front of the user, and running any of them carries no additional fee. Gemma sits in it beside the Llama, Qwen and Mistral families, with entries tuned for specific kinds of work. Local deployments today live in the four-to-nine-billion band.
By default you can. Any open-source model can be added under Settings, and Iternal built it that way because getting a new model approved inside a large enterprise is very difficult. The gate is the build format: models arrive as llama.cpp or OpenVINO builds, and converting one yourself is a command-line process Iternal puts within reach of roughly one percent of customers.
Less than most buyers assume. Open-weight models are now good enough for everyday knowledge work, most users never push past a three-billion-parameter model, and for a mid-sized enterprise a model a couple of revisions behind the frontier usually clears the bar. Agentic work is the exception that raises it.
You still select it: confirm the model is downloaded, then set it in the drop-down. Everything else arrives pre-configured. One caution — the application will let you choose a model your machine cannot carry, because the blocking flag is not in the build yet.
Install as many as you like, but one is active at a time. Each user can be assigned a different model, or switch models on their own device by workflow. Loading one into memory at the start of a chat costs a short delay — another reason to prefer the smallest model that clears your bar.
The restriction is the build format and the size the machine can hold, rather than an allow list Iternal maintains: any OpenVINO-compatible or llama.cpp model can be added. Where a model came from is the sharper question in a regulated enterprise. For more information visit the tuning and provenance page.
Start Small, Measure, Then Move Up
Write the bar down, test the smallest model your fleet can run against your own documents, and stop at the first one that clears it. Every decision that follows has a page above that settles it on its own evidence.