Choosing an AI Tool

How Do You Choose Between
AI Tools and Platforms?

Four axes, the acceptance test that catches a capability gap before purchase, and the request every option answers unchanged.

Built from real buyer questions in our sales meetings

Most AI selections are decided before anyone writes down what winning looks like. A shortlist appears, and the scoring gets invented in the room while the first demonstration runs. Buyers told us people are overwhelmed by the options, and that from a high enough altitude two locally running AI assistants look alike. Criteria written after a demonstration describe it; criteria written before it decide the purchase.

Direct Answer

Score every option on four axes, in this order. Where the data has to sit, whether the product does the specific job you are buying it for, what the license costs once everyone who needs it has it, and whether the purchase consolidates the estate you just rationalized or adds to it. Weight them for your own constraint, then make every option demonstrate your named use case rather than a general capability.

The limit: the demonstration test cuts both ways, including against Iternal. Requiring a demonstration of your named use case is right, and options frequently cannot meet it. We have arrived at evaluations with no demonstration of a prospect’s exact use case and shown a prior one instead, and an hour of live pitch reaches perhaps half of what a product does, so the capability the buyer cared about gets missed. Treat a missing demonstration as information rather than automatic disqualification.

Four things to get in writing before you buy. That the product runs where your material is allowed to sit. That the named workflow has been run end to end on content shaped like yours. The license cost at full headcount and the unit it bills against. Which subscriptions the purchase retires. Iternal answers the targeted questions below plainly: the more advanced document-analysis use cases are not an AirgapAI solution at all.

Comparing options and comparing categories are different exercises. Ranking private AI products against one another is the exercise above; deciding whether the category beats what you already run is the one before it. For more information visit the private AI comparison pillar or the build-or-buy page.

Score Every Option on Four Axes, Weighted for Your Constraint

The four axes below are what buyers argued about in our meetings, ordered by how often a decision turned on them. The weights suit a hard data-placement constraint; move them, and the sheet serves the next business group. Score each option from 1 to 5 before the first demonstration.

Axis Weight A 5 looks like
Where the data has to sit 35% Inference runs inside your own boundary; nothing crosses a border to answer.
Named use-case fit 30% Your exact workflow, run end to end, on content like yours.
Cost at full headcount 20% A license that stops climbing with usage, in a unit you can multiply yourself.
Effect on the estate 15% It retires subscriptions and carries several departmental jobs in one application.

A worked row. One option answers with the network card off: placement scores 5. It showed a prior engagement rather than your contract-review workflow: use-case fit scores 2. It is bought once against the device rather than monthly against every seat: cost scores 5. It retires two subscriptions while adding one application: estate effect scores 4. Weighted, it lands at 3.95 out of 5.

A weighted average ranks options; it does not grant permission. Any axis scoring a 1 is a stop rather than a subtraction. Pilots fail when they lack the right evaluation criteria. Write the sheet first, and the pilot has something to pass.

When the Product Does Not Do the Specific Thing Your Use Case Needs

The objection that ends most evaluations is concrete rather than categorical: the product works, and it leaves open the gap the buyer measured it against. In their own words, document search helps little when simulation testing time is the real constraint, and marketing will not touch a black box its own people cannot edit.

Most of those gaps sit in the design, and design gaps close. Augment the data set with more files and reprocess it, adjust the workflow prompt, or pipe results into the system of record you already run. A model takes correction at any point.

Some gaps are real scope, and those answers beat a yes. Iternal states its own boundaries in the room: the more advanced document-analysis use cases are not an AirgapAI solution, and programming is not the job AirgapAI is built for. An option that names what it will not do has handed you something to plan around.

The acceptance test. Write the use case as one sentence with four parts before anyone demonstrates anything: input, action, output, constraint. “Supplier contracts in PDF, checked for renewal dates, returned as a table citing the source page, on hardware that never leaves our network.” A sentence in that shape is either demonstrated or it is not.

Three questions carry the weight of an evaluation. Put them to every option in writing, in the same words:

Pin it down: questions for your evaluation
  • Will you run our named use case end to end, on content shaped like ours?
    Whether the product does the job you are buying it for, or a job adjacent to it.
  • What does the product not do for this use case today, and what is scheduled rather than shipped?
    The scope boundary in writing, before it becomes a surprise in month three.
  • Where does inference run in the demonstration, and where in production?
    That the data placement you scored is the one you buy.

Score the Shortlist as a Consolidation, Not an Addition

An AI purchase rarely lands in an empty estate. Buyers who had just finished an application rationalization told us a new overlapping tool was unwelcome, that a second solution forces people to remember which tool is allowed where, and that nobody manages four or five separate desktop applications.

The argument that survives that room is consolidation. Apps can be unique use cases; they should not be unique stacks. One application carrying per-department workflows beats five carrying one each, and the rationalization you just ran becomes an argument for the purchase. Iternal builds AirgapAI on that shape: pre-prompted workflow configurations selected by industry or department, prompts and data sets switched on or off per user, positioned as a cost-rationalization play against per-user subscriptions.

Where an incumbent stays, pipe into it. AirgapAI is never a complete replacement for an assistant you already run, and some subscriptions survive. Feed results into the system of record you operate instead of standing a second one beside it.

The Request You Send to Every Option Unchanged

A shortlist is comparable only when every option answered the same request. Send the five lines below without tailoring them; identical requests make different answers mean something.

  • Run our named use case. End to end, in front of us.
  • Use content shaped like ours. Our documents where data-sharing rules allow.
  • Show where the data sat. In the demonstration and in production.
  • Name what the product will not do here. Shipped, scheduled, out of scope.
  • Price it at our headcount. The full license cost and the billing unit.

The fourth line separates the shortlist. An option that answers it plainly has shown you its edges; an option that answers everything confidently has shown you nothing.

Answered elsewhere
FAQ

FAQ: Choosing Between AI Tools

Score each option on four weighted axes: where the data has to sit, whether it does your named use case, the license cost at full headcount, and whether it consolidates the application estate or adds to it. Write the sheet before the first demonstration, and treat any axis scoring a 1 as a stop.

Ask which layer the gap sits in. Most gaps sit in the design and close with a larger data set, a reprocessed corpus or an adjusted workflow prompt. Some are genuine scope: Iternal says openly that the more advanced document-analysis use cases are not an AirgapAI solution.

It does when the purchase is an addition. Score it as a consolidation instead: one application carrying per-department workflows, retiring subscriptions, with prompts and data sets switched on or off per user. Apps can be unique use cases; they should not be unique stacks. Some incumbent subscriptions survive, so pipe results into the system of record.

Run your named use case on both and the differences arrive quickly: where inference actually runs, what the data layer does to accuracy, what unit the license bills against, and what each says it will not do. Iternal differentiates AirgapAI on cost of ownership, hallucination control and security rather than the interface.

Four axes, a weight for each, a 1-to-5 score per option, and a stop rule: data placement, named use-case fit, cost at full headcount, and the effect on the application estate. Weight them for your own constraint, and reuse the sheet by moving the weights.

Write the Sheet Before the First Demonstration

The scorecard costs an afternoon and outlives the shortlist. Name the use case in one sentence, weight the axes, send the same request to every option, and let the replies rank themselves. Bring your sentence to us and we will run it, or tell you plainly it belongs somewhere else.

John Byron Hanby IV
About the Author

John Byron Hanby IV

CEO & Founder, Iternal Technologies

John Byron Hanby IV is the founder and CEO of Iternal Technologies, a leading AI platform and consulting firm. He is the author of The AI Strategy Blueprint and The AI Partner Blueprint, the definitive playbooks for enterprise AI transformation and channel go-to-market. He advises Fortune 500 executives, federal agencies, and the world's largest systems integrators on AI strategy, governance, and deployment.