Turnkey AI API

Bring Your Own Model

Plug-and-play the latest and greatest AI technologies into Iternal's Turnkey AI platform. Connect your preferred LLM - commercial, open-source, or on-premises - and deploy AI capabilities rapidly across your business departments.

TL;DR

Bring Your Own Model, Summarized

Bring Your Own Model (BYOM) is a deployment approach that lets you connect the large language model you prefer — commercial cloud, open-source, or on-premises — to Iternal's Turnkey AI platform, instead of being locked into a single provider's model. The platform provides a common layer for governance, security, and enterprise-grade performance, so you can swap or add models as the technology evolves and deploy AI capabilities rapidly across departments. Supported options include Azure OpenAI, AWS Bedrock, IBM watsonx, Anthropic Claude, open-weight models, and fully private on-premises deployments.

  • Model-agnostic: connect commercial (Azure OpenAI, AWS Bedrock, IBM watsonx, Anthropic Claude), open-source, or on-prem LLMs
  • One platform: unified governance, security, and enterprise performance across every connected model
  • No lock-in: swap or add models as AI advances, without re-platforming
  • Rapid rollout: deploy AI capabilities across sales, marketing, HR, and operations
  • Private path: run on-premises / air-gapped via AirgapAI for regulated data

Ultimate Flexibility

The Bring Your Own Model API enables rapid deployment of various AI technologies across your business departments. Organizations can choose from commercial cloud services, open-source models, or on-premises solutions - all connecting to Iternal's common Turnkey AI platform for unified governance, security, and enterprise-grade performance.

Choose Your LLM

Connect any AI model that fits your organization's needs

Commercial Cloud

Enterprise-ready AI services from leading cloud providers

Azure OpenAI GPT4 AWS Bedrock IBM WatsonX Anthropic Claude

Open Source

Community-driven models with full transparency and control

LLAMA Mistral NVIDIA NIM Falcon

On-Premises

Complete data sovereignty with closed-source solutions

Private Models Air-Gapped Deployment Custom Fine-Tuned
Free download

US Military Deploys Air-Gapped AI

  • Completely air-gapped deployment
  • DoD security requirements met
  • Tactical and strategic applications

Instant download. We'll also email you a copy. No spam.

What Is a Foundation Model, and Which Ones Can You Bring?

A foundation model is a large model pre-trained on broad, unlabeled data that can be adapted to many downstream tasks instead of one. Large language models are the text branch of that family; vision, speech and multimodal models are the others. Any of them with an accessible endpoint or open weights can be connected here.

Foundation model versus task-specific model

A task-specific model is trained on labeled examples for one job — scoring a claim, forecasting demand, classifying a support ticket — and it does that job only. A foundation model is trained once on a broad corpus and then pointed at many jobs through prompting, retrieval or a light adaptation pass. The term was introduced by Stanford's Center for Research on Foundation Models in its 2021 report On the Opportunities and Risks of Foundation Models; the EU AI Act describes the same class of system as a general-purpose AI model. The practical difference for a buyer is procurement: you acquire one capable model and adapt it across departments, rather than commissioning a new model for every use case.

Where large language models sit inside the family

Every large language model is a foundation model; not every foundation model is a large language model. The family splits by what the model was pre-trained to read and produce: language models work on text, vision models on images and video, speech models on audio, and multimodal models across two or more of those at once. Most enterprise deployments start with a language model because documents and conversations are where the work is, then add a speech model for transcription and a vision model for drawings, scans and inspection imagery.

The three ways a foundation model arrives

Commercial models reach you as a hosted endpoint with the weights held by the provider. Open-weight models are downloaded under a license and run wherever you choose, including fully disconnected hardware. Private models are your own copy — an open-weight model you have adapted, or a closed model licensed for on-premises use. Iternal's Turnkey AI platform accepts all three through the same connection layer, which is what makes switching models a configuration change rather than a rebuild.

Foundation models enterprises connect, by family

Model Family Provider First released Weights & license How it connects
GPT-4o Text + vision OpenAI May 2024 Proprietary, hosted API Azure OpenAI or OpenAI endpoint
Claude Sonnet 4 Language Anthropic May 2025 Proprietary, hosted API Anthropic or AWS Bedrock endpoint
Llama 3.1 405B Language Meta July 2024 Open weights, Llama 3.1 Community License Self-hosted, NVIDIA NIM or AWS Bedrock
Mistral 7B Language Mistral AI September 2023 Open weights, Apache 2.0 Self-hosted or on-device
Granite 3.0 Language IBM October 2024 Open weights, Apache 2.0 IBM watsonx or self-hosted
Phi-4 Language (small) Microsoft December 2024 Open weights, MIT Self-hosted or on-device
Qwen2.5 (7B–32B) Language Alibaba September 2024 Open weights, Apache 2.0 Self-hosted or on-device
gpt-oss-20b Language OpenAI August 2025 Open weights, Apache 2.0 Self-hosted or on-device
Whisper large-v3 Speech OpenAI November 2023 Open weights, MIT Self-hosted transcription service
Segment Anything 2 Vision Meta July 2024 Open weights, Apache 2.0 Self-hosted vision service

Dates are the first public release of the named version. License terms differ by model size and change between releases, so confirm the model card and its acceptable-use policy before deployment. Which of these fits a given workload is a separate question, worked through on the LLM selection guide.

Bringing an open-weight foundation model on-premises

Open weights are what make a private deployment possible: the model file sits on your hardware, so inference never leaves the boundary and there is no per-token bill for usage. AirgapAI runs open-weight models directly on an AI PC or a local server with no network connection at all, and Blockify structures your source documents so the connected model answers from governed content instead of raw files. For the infrastructure side of the same decision — sizing, quantization, serving and hardware — see how to deploy an LLM on-premise.

How Enterprises Adapt a Foundation Model

Four adaptation paths exist, in ascending cost: prompting, retrieval over your own documents, fine-tuning the weights, and pre-training a model from scratch. Most enterprise work lands on the first two, because they change what the model knows without changing the model, and both survive a swap to a newer foundation model.

Prompting and context

Instructions, examples and system context shape the output without touching weights. Fastest to change, and portable across every connected model.

Retrieval over your data

The model is handed passages from your own documents at query time, so answers cite current source material the model was never trained on.

Fine-tuning

A training pass adjusts the weights for a format, a tone or a narrow domain. It buys consistency, not fresh knowledge, and it is re-run when you change models.

Pre-training from scratch

Building a new foundation model is a research program with a capital budget to match, and it is the right answer for very few organizations.

The full decision is documented elsewhere on this site rather than repeated here: the trade-offs between the four paths are worked through in Custom AI Models: Train vs Fine-Tune vs RAG vs Prompt Engineering, and the two most common options are compared directly on RAG vs fine-tuning. For picking the model itself, start with the LLM selection guide.

Enterprise Benefits

Why organizations choose Bring Your Own Model

Flexibility

Switch models without changing your applications

Security

Unified governance across all AI models

Rapid Deployment

Deploy AI capabilities in weeks, not months

IdeaBlocks Power

78X more accurate AI with governed data

Future-Proof

Adopt new models as they become available

Provider Choice

Avoid lock-in with any single provider

Cost Control

Choose models that match your budget

Team Enablement

Deploy AI across all departments

Getting Started

1

Subscribe

Start your annual Turnkey AI subscription to access the platform

2

Connect

Configure your preferred LLM connection(s) with our integration support

3

Deploy

Go live with AI capabilities across your organization

Integration Timeline: The Bring Your Own Model API requires a 1-2 week service integration period. Our team works with you to configure, test, and validate your LLM connection(s). Customers manage deployment and maintenance of their chosen LLM connection(s).

Master AI Skills with Iternal AI Academy

Elevate your team's AI capabilities with comprehensive training. Learn to leverage AI models effectively and get the most from your Turnkey AI deployment.

810+
Courses
50+
Industries
200+
Job Roles
  • LLM integration and API courses
  • Hands-on exercises with immediate application
  • 2-4 hour courses in 10-minute lessons
  • New content added weekly
Start Learning Today

Unlimited Access

All courses, all industries, all skill levels

$199/year

Bring Your Own Foundation Model: Frequently Asked Questions

Can I bring my own foundation model to Iternal’s Turnkey AI platform?

Yes. The Bring Your Own Model API connects a commercial endpoint such as Azure OpenAI, AWS Bedrock, IBM watsonx or Anthropic Claude, an open-weight model you host yourself, or a private on-premises deployment. Integration takes one to two weeks, and the model you connect stays under your contract with its provider.

What is the difference between a foundation model and an LLM?

A foundation model is any model pre-trained on broad data and adaptable to many tasks. A large language model is the text branch of that family. Vision, speech and multimodal models are foundation models too, so every LLM is a foundation model but not every foundation model is an LLM.

Which open-weight foundation models can run on-premises?

Models published with downloadable weights can run on your own hardware: the Llama, Mistral, Qwen, Granite, Phi and gpt-oss language families, Whisper for speech, and Segment Anything for vision. What limits you in practice is memory and accelerator capacity rather than the license, so parameter count and quantization decide which models fit a given server or AI PC. The LLM parameter size guide covers that sizing question.

Do I have to fine-tune a foundation model to use my own data?

No. Retrieval hands the model passages from your documents at query time, which keeps answers current and traceable without a training run. Fine-tuning is for consistency of format, tone or a narrow domain, and it has to be repeated whenever you change models. RAG vs fine-tuning works through the trade-off.

How do I decide which foundation model to connect?

Decide on the constraints first: where inference is allowed to run, the context length the work needs, latency and cost per token, and the license terms your legal team will accept. Model-by-model selection is covered on the LLM selection guide, which is kept current as new releases land.

Can a foundation model run in an air-gapped environment?

Yes, if its weights are downloadable. An open-weight model, its runtime and its data set can be installed on a machine with no network connection, which is how AirgapAI runs on an AI PC or a local server for classified, controlled or otherwise restricted material.

What licenses apply to open-weight foundation models?

They vary. Some releases are Apache 2.0 or MIT, which place few restrictions on commercial use. Others use a community or research license with acceptable-use terms, user thresholds or naming requirements attached. Licenses also differ between sizes in the same family, so check the model card for the exact weights you plan to deploy.

Bring Your AI Vision to Life

Connect your preferred LLM to Iternal's enterprise-ready platform and deploy AI capabilities across your organization in weeks, not months. Bringing your own model is one piece of the full private LLM suite — see the enterprise hub for the complete private, local, and air-gapped stack.