Bring Your Own Model
Plug-and-play the latest and greatest AI technologies into Iternal's Turnkey AI platform. Connect your preferred LLM - commercial, open-source, or on-premises - and deploy AI capabilities rapidly across your business departments.
Bring Your Own Model, Summarized
Bring Your Own Model (BYOM) is a deployment approach that lets you connect the large language model you prefer — commercial cloud, open-source, or on-premises — to Iternal's Turnkey AI platform, instead of being locked into a single provider's model. The platform provides a common layer for governance, security, and enterprise-grade performance, so you can swap or add models as the technology evolves and deploy AI capabilities rapidly across departments. Supported options include Azure OpenAI, AWS Bedrock, IBM watsonx, Anthropic Claude, open-weight models, and fully private on-premises deployments.
- Model-agnostic: connect commercial (Azure OpenAI, AWS Bedrock, IBM watsonx, Anthropic Claude), open-source, or on-prem LLMs
- One platform: unified governance, security, and enterprise performance across every connected model
- No lock-in: swap or add models as AI advances, without re-platforming
- Rapid rollout: deploy AI capabilities across sales, marketing, HR, and operations
- Private path: run on-premises / air-gapped via AirgapAI for regulated data
Ultimate Flexibility
The Bring Your Own Model API enables rapid deployment of various AI technologies across your business departments. Organizations can choose from commercial cloud services, open-source models, or on-premises solutions - all connecting to Iternal's common Turnkey AI platform for unified governance, security, and enterprise-grade performance.
US Military Deploys Air-Gapped AI
- Completely air-gapped deployment
- DoD security requirements met
- Tactical and strategic applications
Instant download. We'll also email you a copy. No spam.
What Is a Foundation Model, and Which Ones Can You Bring?
A foundation model is a large model pre-trained on broad, unlabeled data that can be adapted to many downstream tasks instead of one. Large language models are the text branch of that family; vision, speech and multimodal models are the others. Any of them with an accessible endpoint or open weights can be connected here.
Foundation model versus task-specific model
A task-specific model is trained on labeled examples for one job — scoring a claim, forecasting demand, classifying a support ticket — and it does that job only. A foundation model is trained once on a broad corpus and then pointed at many jobs through prompting, retrieval or a light adaptation pass. The term was introduced by Stanford's Center for Research on Foundation Models in its 2021 report On the Opportunities and Risks of Foundation Models; the EU AI Act describes the same class of system as a general-purpose AI model. The practical difference for a buyer is procurement: you acquire one capable model and adapt it across departments, rather than commissioning a new model for every use case.
Where large language models sit inside the family
Every large language model is a foundation model; not every foundation model is a large language model. The family splits by what the model was pre-trained to read and produce: language models work on text, vision models on images and video, speech models on audio, and multimodal models across two or more of those at once. Most enterprise deployments start with a language model because documents and conversations are where the work is, then add a speech model for transcription and a vision model for drawings, scans and inspection imagery.
The three ways a foundation model arrives
Commercial models reach you as a hosted endpoint with the weights held by the provider. Open-weight models are downloaded under a license and run wherever you choose, including fully disconnected hardware. Private models are your own copy — an open-weight model you have adapted, or a closed model licensed for on-premises use. Iternal's Turnkey AI platform accepts all three through the same connection layer, which is what makes switching models a configuration change rather than a rebuild.
Foundation models enterprises connect, by family
| Model | Family | Provider | First released | Weights & license | How it connects |
|---|---|---|---|---|---|
| GPT-4o | Text + vision | OpenAI | May 2024 | Proprietary, hosted API | Azure OpenAI or OpenAI endpoint |
| Claude Sonnet 4 | Language | Anthropic | May 2025 | Proprietary, hosted API | Anthropic or AWS Bedrock endpoint |
| Llama 3.1 405B | Language | Meta | July 2024 | Open weights, Llama 3.1 Community License | Self-hosted, NVIDIA NIM or AWS Bedrock |
| Mistral 7B | Language | Mistral AI | September 2023 | Open weights, Apache 2.0 | Self-hosted or on-device |
| Granite 3.0 | Language | IBM | October 2024 | Open weights, Apache 2.0 | IBM watsonx or self-hosted |
| Phi-4 | Language (small) | Microsoft | December 2024 | Open weights, MIT | Self-hosted or on-device |
| Qwen2.5 (7B–32B) | Language | Alibaba | September 2024 | Open weights, Apache 2.0 | Self-hosted or on-device |
| gpt-oss-20b | Language | OpenAI | August 2025 | Open weights, Apache 2.0 | Self-hosted or on-device |
| Whisper large-v3 | Speech | OpenAI | November 2023 | Open weights, MIT | Self-hosted transcription service |
| Segment Anything 2 | Vision | Meta | July 2024 | Open weights, Apache 2.0 | Self-hosted vision service |
Dates are the first public release of the named version. License terms differ by model size and change between releases, so confirm the model card and its acceptable-use policy before deployment. Which of these fits a given workload is a separate question, worked through on the LLM selection guide.
Bringing an open-weight foundation model on-premises
Open weights are what make a private deployment possible: the model file sits on your hardware, so inference never leaves the boundary and there is no per-token bill for usage. AirgapAI runs open-weight models directly on an AI PC or a local server with no network connection at all, and Blockify structures your source documents so the connected model answers from governed content instead of raw files. For the infrastructure side of the same decision — sizing, quantization, serving and hardware — see how to deploy an LLM on-premise.
How Enterprises Adapt a Foundation Model
Four adaptation paths exist, in ascending cost: prompting, retrieval over your own documents, fine-tuning the weights, and pre-training a model from scratch. Most enterprise work lands on the first two, because they change what the model knows without changing the model, and both survive a swap to a newer foundation model.
Prompting and context
Instructions, examples and system context shape the output without touching weights. Fastest to change, and portable across every connected model.
Retrieval over your data
The model is handed passages from your own documents at query time, so answers cite current source material the model was never trained on.
Fine-tuning
A training pass adjusts the weights for a format, a tone or a narrow domain. It buys consistency, not fresh knowledge, and it is re-run when you change models.
Pre-training from scratch
Building a new foundation model is a research program with a capital budget to match, and it is the right answer for very few organizations.
The full decision is documented elsewhere on this site rather than repeated here: the trade-offs between the four paths are worked through in Custom AI Models: Train vs Fine-Tune vs RAG vs Prompt Engineering, and the two most common options are compared directly on RAG vs fine-tuning. For picking the model itself, start with the LLM selection guide.
Enterprise Benefits
Why organizations choose Bring Your Own Model
Flexibility
Switch models without changing your applications
Security
Unified governance across all AI models
Rapid Deployment
Deploy AI capabilities in weeks, not months
IdeaBlocks Power
78X more accurate AI with governed data
Future-Proof
Adopt new models as they become available
Provider Choice
Avoid lock-in with any single provider
Cost Control
Choose models that match your budget
Team Enablement
Deploy AI across all departments
Getting Started
Subscribe
Start your annual Turnkey AI subscription to access the platform
Connect
Configure your preferred LLM connection(s) with our integration support
Deploy
Go live with AI capabilities across your organization
Integration Timeline: The Bring Your Own Model API requires a 1-2 week service integration period. Our team works with you to configure, test, and validate your LLM connection(s). Customers manage deployment and maintenance of their chosen LLM connection(s).
Master AI Skills with Iternal AI Academy
Elevate your team's AI capabilities with comprehensive training. Learn to leverage AI models effectively and get the most from your Turnkey AI deployment.
- LLM integration and API courses
- Hands-on exercises with immediate application
- 2-4 hour courses in 10-minute lessons
- New content added weekly
Unlimited Access
All courses, all industries, all skill levels
Bring Your Own Foundation Model: Frequently Asked Questions
Can I bring my own foundation model to Iternal’s Turnkey AI platform?
Yes. The Bring Your Own Model API connects a commercial endpoint such as Azure OpenAI, AWS Bedrock, IBM watsonx or Anthropic Claude, an open-weight model you host yourself, or a private on-premises deployment. Integration takes one to two weeks, and the model you connect stays under your contract with its provider.
What is the difference between a foundation model and an LLM?
A foundation model is any model pre-trained on broad data and adaptable to many tasks. A large language model is the text branch of that family. Vision, speech and multimodal models are foundation models too, so every LLM is a foundation model but not every foundation model is an LLM.
Which open-weight foundation models can run on-premises?
Models published with downloadable weights can run on your own hardware: the Llama, Mistral, Qwen, Granite, Phi and gpt-oss language families, Whisper for speech, and Segment Anything for vision. What limits you in practice is memory and accelerator capacity rather than the license, so parameter count and quantization decide which models fit a given server or AI PC. The LLM parameter size guide covers that sizing question.
Do I have to fine-tune a foundation model to use my own data?
No. Retrieval hands the model passages from your documents at query time, which keeps answers current and traceable without a training run. Fine-tuning is for consistency of format, tone or a narrow domain, and it has to be repeated whenever you change models. RAG vs fine-tuning works through the trade-off.
How do I decide which foundation model to connect?
Decide on the constraints first: where inference is allowed to run, the context length the work needs, latency and cost per token, and the license terms your legal team will accept. Model-by-model selection is covered on the LLM selection guide, which is kept current as new releases land.
Can a foundation model run in an air-gapped environment?
Yes, if its weights are downloadable. An open-weight model, its runtime and its data set can be installed on a machine with no network connection, which is how AirgapAI runs on an AI PC or a local server for classified, controlled or otherwise restricted material.
What licenses apply to open-weight foundation models?
They vary. Some releases are Apache 2.0 or MIT, which place few restrictions on commercial use. Others use a community or research license with acceptable-use terms, user thresholds or naming requirements attached. Licenses also differ between sizes in the same family, so check the model card for the exact weights you plan to deploy.
Bring Your AI Vision to Life
Connect your preferred LLM to Iternal's enterprise-ready platform and deploy AI capabilities across your organization in weeks, not months. Bringing your own model is one piece of the full private LLM suite — see the enterprise hub for the complete private, local, and air-gapped stack.