Choosing Where AI Runs

Should AI Run on the Laptop, on Your Own Server,
or in a Hosted Cloud?

Two questions settle the placement: who may see the material, and how many people need it at once. What each placement buys you, what it costs, and the hybrid most deployments land on.

Built from real buyer questions in our sales meetings

Placement looks like plumbing and behaves like strategy. Install the software on a laptop and one person gets a private assistant whose every answer costs nothing extra; install it on a server and a department shares one curated corpus and a bigger model. Buyers asked us both halves in the same conversation, again and again — can we host it ourselves, and does it need a server behind it?

Direct Answer

Two questions choose the placement: who may see the material, and how many people need it at once. Run it on the device when one person works on sensitive material and you want inference that adds nothing per query. Run it on a server you own, in your network or your own cloud tenant, when several people need one shared corpus or a model larger than a laptop holds. Choose a hosted deployment when neither constraint binds. A hybrid is normal: the assistant on the device, the larger model on a server behind it.

The limit sits on the device side. AirgapAI on a laptop is single-tenanted by design, which is what satisfies a rule that nothing may leave the machine — and it means the device deployment does not centralize a corpus across users, has no shared multi-user instance (people must sit in the same Windows profile to share one interface and chat history), and carries no collaboration on shared documents or sessions. The server deployment fixes all three and takes materially more effort and time than the one-click client install. Choosing the device is choosing single-user working. Choosing the server is choosing a deployment project.

Get the placement written down before you buy. Establish which component sits where, what a device does when the server behind it is unreachable, and how long a move from a hosted pilot to your own hardware takes. Iternal answers each in writing — the evaluation questions below are the short list.

Placement is not permission. Whether your material may leave the boundary that governs it is decided elsewhere. For more information visit the running AI on data that cannot leave section, the offline and air-gapped page for what keeps working once the network is gone, and the sizing page for what the placement has to be specified as. Placement answers one question only: where the software runs.

Three Placements, and the Two Questions That Choose Between Them

The first question is permission: how sensitive is the material, and whose infrastructure may touch it. The second is scale: how many people need the same answers from the same corpus at once. Model size rides along with the second, because the machine that serves many people is also the machine that holds a larger model.

Placement Choose it when The trade you are making Default
On each user's device One person works on material that must not leave their machine, and each user needs answers from their own data sets. Single-user working: no corpus centralized across users, no shared session or history, and the model has to fit the hardware. Start here: a one-click installer, nothing leaving the machine, unlimited use, no token cost.
On a server you own Several people need one curated corpus, or the work needs a model bigger than a laptop holds. A deployment project: someone stands up and runs the model, which is more technical than the client installer. A relay architecture with the desktop app as a thin client, and all processing happening on the server.
Hosted in a boundary you control Neither constraint binds, or you need working capacity before your own hardware arrives. The material sits in a cloud boundary, so permission has to be settled before placement. Containers inside your own tenant: Iternal delivers on-premises or into the customer tenant, not as a shared service.

When the Requirement Is Where the Workload Sits

Some buyers arrive with the placement already decided, and say so in their own words: the model has to sit on our own servers, because clients ask about the security of their data; it has to be on-premises because it is defense work; law firms will not send their most sensitive information to the cloud. Each sentence names a boundary, and each boundary names a different placement:

  • It cannot leave the device. The device placement. AirgapAI content is single-tenanted and stored locally, so one user never sees another user's work.
  • It cannot leave our servers. A server inside your own network. The models and data sets live there, all processing happens there, and the app is a thin client.
  • It cannot leave our control boundary. Your own tenant. Iternal states that AirgapAI can be deployed in a virtual private cloud inside the customer control boundary with role-based access, delivered as containers.
  • It cannot touch a network at all. A sealed placement, supported today: AirgapAI is deployed in SCIF facilities, and Iternal built the technology for federal and classified work.

What qualifies a deployment as genuinely sealed is a separate standard, with its own test. For more information visit the offline and air-gapped page.

The Hybrid: A Local Assistant Pointed at a Bigger Model

Most estates stop treating the choice as either-or once they see the middle option. AirgapAI connects only to localhost by design, and a light relay on the device points that same client at a model hosted elsewhere — a server inside your intranet, a workstation-class box under a desk, or a trusted endpoint in a cloud you control. The application stays where the user is. The compute moves to where the capacity is.

What the hybrid changes. Models and data sets move onto the server, all processing happens there, and one common set of data sets is mapped down to every device instead of being rebuilt per laptop. A department gets the very large data sets and the models a laptop could not hold. Iternal states that running the models on a server needs no additional model licensing beyond one AirgapAI instance per user, and that more customers are moving to the server-hosted arrangement.

What the hybrid changes when the network is unavailable. The dependency moves with the model. With the model on the device, a dead network changes nothing about answering. With the model on a server, the local network becomes load-bearing: lose it and the user falls back to whatever their own machine holds. Keeping a local model on each device alongside the server pointer is a different posture from a thin client with nothing behind it. Choose deliberately.

For more information visit the hybrid AI architecture page, which sets out the same device-server-hosted split as a design pattern across an estate, or the on-premise AI chat guide, which works the same placement decision through for a chat assistant a whole team shares.

Start Hosted, Move On-Premises When the Hardware Lands

Hardware arrives on procurement time, and evaluations run on business time. The path Iternal recommends for buyers caught between the two is to start in the cloud to prove the deployment in operation, then buy the servers on-premises for security, ownership and future-proofing. The cloud deployment is not throwaway work: the same stack moves onto the on-premises server once it arrives, a cutover Iternal describes as measured in days rather than a re-implementation.

The same staging logic applies to the material: a pilot can begin on public or unclassified content and take on classified material later, once the placement that material requires exists.

The Debate Most Buyers Have Not Settled

Plenty of organizations have never made the call at all, and described the drift plainly: every conversation defaults to cloud; internal offline work was paused and lost traction as attention moved to cloud AI; the team that liked the on-premises proposal ran its proof of concept on a hyperscaler anyway. Others want the opposite and doubt they can have it: enterprises that want AI on their core data and intellectual property, unsure they can run it internally.

Underneath the drift sits a quality doubt. Buyers measure an on-device product against what they can swipe a credit card and get from a frontier model today, and assume a private deployment must be the weaker option. Placement does not set answer quality, though. The model you point the client at and the data you feed it do. AirgapAI ships with an open Meta Llama model in the 8-billion-parameter class, runs other open models, and the same client can be pointed at a much larger model on your own server without the material moving anywhere new. Buyers who love the device and still want large models in the data center are describing the hybrid.

Where ambiguity is worth money — timing, fallback behavior, who holds the keys — put it in writing:

Pin it down: questions for your evaluation
  • Which components sit on the device and which sit on the server in the configuration you are quoting us?
    The placement of every moving part, not just the assistant.
  • What does a user device do when the server or the relay behind it is unreachable?
    Whether your people keep working through a network outage or stop.
  • How long does the move from a hosted pilot to our own on-premises server take, and what is the cutover plan?
    The migration timing for your configuration, in writing.
  • Is our deployment delivered into our own tenant rather than a shared service, and who holds the administrative keys?
    That a hosted deployment stays inside your control boundary, with access you administer.
Answered elsewhere
FAQ

FAQ: Choosing Where AI Runs

Either. AirgapAI runs the model on the local workstation with no server required, which is how most deployments start. It can also act as a thin client against a model on a server inside your network or in a cloud boundary you control, reached through a relay on the device, and then all processing happens on the server. Put the model where the constraint is.

Yes to both. Iternal states that AirgapAI can be deployed on-premises, or in a virtual private cloud inside your own control boundary with role-based access, delivered as containers into your tenant rather than run for you as a shared service. What the deployment asks of you is an owner for the server, the model and the access rules.

Yes, through a relay. AirgapAI connects only to localhost by design, so a light relay on the device points it at a larger model on your own server or another inference endpoint you run. Iternal states that no additional model licensing is required beyond one AirgapAI instance per user. The cost: you deploy and operate that model yourself, which is more technical than the client installer.

Not on the device deployment. AirgapAI on a laptop has no shared multi-user instance — people have to sit in the same Windows profile to share one interface and chat history — it does not centralize a corpus across users, and it carries no collaboration on shared documents or sessions. The server deployment answers all three: models and data sets live on the server, mapped down to each device. Expect materially more effort than the client install.

That is the path Iternal recommends when hardware is on order: prove the deployment in operation in the cloud, then buy the servers on-premises for security, ownership and future-proofing. The cloud work is not thrown away — the same stack moves onto the on-premises server once it arrives, a cutover Iternal describes as measured in days rather than a rebuild. Ask for the elapsed time and the cutover plan for your configuration in writing.

Match the placement to the wording of the rule. Cannot leave the machine: run it on the device, where AirgapAI content is single-tenanted and stored locally. Cannot leave your servers: put the models and data sets on a server inside your network and use the app as a thin client. Cannot leave your control boundary: your own tenant or a virtual private cloud satisfies the rule. Can never touch a network: the sealed placement is supported today, and AirgapAI is deployed in SCIF facilities.

Decide the Placement Before You Decide Anything Else

Write down two sentences before the next infrastructure meeting: who may see this material, and how many people need it at once. They choose the row in the table above, and every later decision inherits from that row. Install the client on one machine first — the cheapest way to learn which row you are in.

John Byron Hanby IV
About the Author

John Byron Hanby IV

CEO & Founder, Iternal Technologies

John Byron Hanby IV is the founder and CEO of Iternal Technologies, a leading AI platform and consulting firm. He is the author of The AI Strategy Blueprint and The AI Partner Blueprint, the definitive playbooks for enterprise AI transformation and channel go-to-market. He advises Fortune 500 executives, federal agencies, and the world's largest systems integrators on AI strategy, governance, and deployment.