Cutting Through AI Hype

How Much of the AI Pitch Is Hype,
and How Do Buyers Get Taken Advantage Of?

Three questions that separate a measured claim from a marketing one, applied to Iternal’s own numbers and to the statistic the whole market repeats.

Built from real buyer questions in our sales meetings

Every AI number you are shown was measured somewhere, by someone, under conditions somebody chose. Whether anyone will tell you what they were is the whole test. Buyers put it less politely: the market is overhyping this, and half of what people believe about AI is wrong.

Direct Answer

The hype is real; the defense is procedural. Nobody out-argues a confident seller. One rule handles it, applied to every figure including ours: a number arrives with the conditions it was taken under, or it does not count. Three questions do the job. How was it measured? Can we reproduce it here? What would prove it false?

The limit: run that test on Iternal and it finds something. Iternal has no single clean tagline, three people inside the company would each answer what it does differently, and one partner executive needed two years to grasp the value proposition. That disqualifies nobody. It is exactly the condition in which to demand a measured number rather than a description, from us as much as from anyone else.

What to get in writing before money moves. Require the conditions behind every figure: hardware, data set, comparison. Require the right to re-run the test on your own machines with your own documents. Require the seller to name the result that would prove the claim false. A claim nothing could disprove is not a claim; it is copy.

A figure that fails those questions is not evidence, whichever way it points. The line that 95 percent of AI efforts fail travels further than any product claim here, and Iternal has run it in its own advertising. It dies at question one. For more information on a product figure published with its condition attached, visit the answers you can trust page.

What the Market Looks Like From the Buying Side

Buyers described the same discomfort repeatedly, rarely politely. They are a little scared of being taken advantage of because they are new to this. They ask firms to scope the work and get back either lowball numbers or inflated dollars that make no sense beside each other. They watch peers blindly trust whatever a frontier model produces. One executive named the trap in claiming expertise here: it leads you down a fast rabbit hole. The other half of the problem is belief rather than deception — that AI demands constant cloud connectivity, which it does not, or that only young people can adopt it, which is backwards.

Three Questions That Separate a Measured Claim From a Marketing One

A buyer cannot audit a technology. A buyer can audit a sentence. Run every headline figure through the same three questions and the weak ones fall out.

Ask this A claim that survives A claim that does not
How was it measured? Names the silicon, the data set and the comparison. A bare percentage nobody in the room can source.
Can we reproduce it here? Comes with a way to re-run it on your hardware and documents. Reproducible only on the seller’s machine and corpus.
What would prove it false? Names the condition under which the number drops, and shows the drop. Shaped so no result could disprove it — or confirm it.

Question two carries extra weight in AI. Benchmarks are not standardized, and throughput belongs to the exact hardware and model that produced it, so extrapolating from someone else’s run is guesswork wearing a decimal point.

Turning the Test on Our Own Numbers

An audit rule that exempts the company publishing it is worth nothing, so Iternal’s own figures go first.

How it was measured, published with the figure. Iternal states a large accuracy gain for pre-processing a corpus with Blockify, and states in the same breath that the gain depends on how duplicated the corpus was, with a much smaller multiple on data already cleaned. Read the figure where its condition sits beside it, on the answers you can trust page.

Reproducibility, built in. AirgapAI benchmarks your own machine at setup and whenever a model is added, reporting tokens per second and prefill time in settings and in an exportable report. Iternal publishes the conditions behind its own runs too: the Blockify server benchmark ran on a Google Cloud Xeon 6 instance, the only provider offering that silicon at the time. Where the conditions were thin they were thin out loud — one run covered a single server specification, because only one was available.

What would prove it false is yours to run. Which is why the test belongs on your own material, not a prepared demonstration set:

Pin it down: questions for your evaluation
  • Which data set produced this figure, and what condition was that data in before the run?
    Whether the number describes material like yours or an ideal corpus.
  • Can we re-run the benchmark on our own hardware and documents during the evaluation?
    Whether the result transfers onto the fleet you actually own.
  • What result in our test would count as this claim failing, and how would we see it?
    That the claim is falsifiable, which separates a measurement from a slogan.

When the Service Itself Starts Erroring Out

Claims decay a second way that no benchmark catches: the product that tested well stops behaving. Buyers were specific. A hosted assistant that works beautifully for an hour and degrades on the next, as though quietly quantized once traffic climbs. A search over processed documents returning a rate-limit error instead of content. Underneath both, one doubt — was any of this tested before it shipped?

Reliability belongs on a page about hype because it is where a claim gets settled after the sale. Iternal holds builds before they reach everyone and is moving to named releases with published notes rather than a stream of small changes. The counterpart: Iternal runs no separate test department, so release discipline is a practice rather than a gate — fair to probe, here as anywhere.

Pin it down: questions for your evaluation
  • Are changes packaged as named releases with published notes, or shipped as a continuous stream?
    Whether you can pin a version and know what moved.
  • If a release breaks something here, what is the documented path back to the previous build?
    Recovery time agreed before the incident rather than during it.

The Number Everyone Repeats, Put Through the Same Test

One figure has done more work in AI selling than any product claim: that 95 percent of AI efforts fail. It opens keynotes and justifies budgets. Iternal has advertised it. It is also the cleanest example of a number that dies at question one.

The subject moves between tellings. The figure covers pilots, then projects, then implementations, then initiatives, then investments — five populations with five denominators. A pilot that ends without a rollout is an ordinary result of experimenting; a failed investment is a write-off.

The source comes and goes. An MIT study is named on some occasions and left out on others, and the figure circulates happily without it. A number that survives losing its citation has stopped working as evidence.

Failure is never defined. Canceled, delayed, over budget, still running without a production system — nobody says which. Question three has no answer, so no result could disprove the claim.

AI efforts do fail; buyers put the cause at staffing and missing evaluation criteria, not technology. The point is narrower: a figure you cannot audit should not move your budget in either direction, including Iternal’s own.

Answered elsewhere
FAQ

FAQ: Hype, Claims and What Is Actually Measurable

Three questions, in order. How was it measured — on which hardware, against which data set, compared with what? Can we reproduce it here during the evaluation? What result would prove it false? A figure that survives all three is evidence; one that fails the first is a slogan with a decimal point.

Only loosely. Benchmarks are not standardized, and a higher score does not reliably mean a difference a user would notice. Throughput belongs to the hardware and model that produced it, so extrapolation is estimation. Treat a published score as a reason to test. AirgapAI benchmarks the machine it installs on, so your number is your own.

Treat it as market rhetoric rather than measurement. Its subject drifts between tellings — pilots, projects, implementations, initiatives, investments — which are different populations with different denominators. An MIT study is named on some occasions and left out on others, and nobody defines failing. Iternal has advertised the line, and it still fails the measurement question.

Ask about release practice. Are changes packaged as named releases with published notes, or shipped as a continuous stream? What happens to a build between development and your fleet, and who signs it off? What is the path back when a release breaks something? Iternal holds builds before general release and runs no separate test department.

Backwards. Experienced people know what good work looks like, which is the judgment needed to tell high-quality AI output from low. The colleague most likely to catch a fabricated citation has written hundreds of the real thing. The adoption variable is willingness to check the work, not age.

Start With the Number You Produced Yourself

The fastest way out of an argument about claims is a measurement you took. Install AirgapAI on one machine, let it benchmark the hardware, load your own documents. That afternoon outranks every figure here, ours included.

John Byron Hanby IV
About the Author

John Byron Hanby IV

CEO & Founder, Iternal Technologies

John Byron Hanby IV is the founder and CEO of Iternal Technologies, a leading AI platform and consulting firm. He is the author of The AI Strategy Blueprint and The AI Partner Blueprint, the definitive playbooks for enterprise AI transformation and channel go-to-market. He advises Fortune 500 executives, federal agencies, and the world's largest systems integrators on AI strategy, governance, and deployment.