Choosing Your First AI Project

How Do You Identify Which AI Use Cases Are Yours,
and Choose the First to Pilot?

A scoring sheet you can fill in this week, the three filters that separate a pilot from a demo, and the minimum definition of a proof worth running.

Built from real buyer questions in our sales meetings

Any business can list a hundred things AI might do. Almost none can name the one to do first. The list is never the constraint; the ranking is — and buyers described the symptom to us the same way over and over: a discovery call that went everywhere and produced no concrete list, an AI leader inundated with unqualified people bringing ideas, pilots that were solutions looking for a problem. Two questions sit underneath all of it. Which use cases are actually ours? And which one do we prove first?

Direct Answer

Work backwards from repetitive hours and expensive errors. Any process with a document in, a judgment in the middle and a document out is a candidate. Score every candidate on three axes — impact, ease of implementation, and the technical debt it leaves behind — then apply three hard filters to the top of the scored list: you already own the data, you can state what the process costs today in hours, and one named person will accept the output as their responsibility. Scoring and choosing are the same exercise run twice. Anything missing one of those three filters becomes a demo that never ships.

The limit sits inside the shortlist itself. A candidate can pass all three filters and still be the first of its kind. Iternal has taken on work it had not delivered for anyone before — document remediation for higher education began exactly that way — and a candidate can score well precisely because nobody has done it yet. Novelty is a premium, not a discount. Scoring tells you which candidate is worth the risk; it tells you nothing about whether anyone has shipped it. Ask that as its own question and price the answer.

Write the success threshold down before the work starts. Iternal groups first-pass criteria into accuracy, usability and throughput; set your numbers in those terms, name who signs that they were met, and run the proof on your own documents wherever your data-sharing policy allows. Mock material proves the flow; your material proves the result. Treat the claim that 95% of AI projects fail as a marketing line rather than a reason to act — its subject changes every time it is repeated, and no buyer has ever tied it, in front of us, to a budget that stalled.

Selection and cataloging are different jobs. Selection is the work set out below: finding candidates, scoring them, and picking the first to prove. For more information on what other organizations already run, visit the use case catalog. For more information on the guided interview that runs the same exercise as software, visit the AI Blueprint Builder page.

Find the Candidates First, Then Rank Them on Three Axes

Repetition is the tell. The work worth automating first is work somebody already does the same way every week, on documents that look alike, with a judgment in the middle that a competent colleague could describe in one sentence. Four shapes recur across Iternal engagements often enough to serve as a first sweep: enhancing content that already exists, generating new content from it, working through long documents nobody has time to read, and answering the same quick questions over and over.

Walk each department against those four shapes and candidates arrive faster than any brainstorm produces them. Read them for bias: lists gathered from infrastructure teams lean toward automation more than AI. A scheduled script is a fine outcome, and it belongs in a different budget line.

Ranking is where the exercise earns its keep. Iternal stack-ranks candidates by impact, by ease of implementation and by the technical debt each one creates, and a candidate that lands at the bottom is telling you something concrete: heavy technical debt, or a high cost to implement. The top of that ranked list is where the three hard filters go to work. One sheet carries both halves, and the filter rows are the ones that decide:

Column on the sheet What you write in it What a weak entry looks like
Candidate One process, named in the words of the team that runs it. A department name with the word AI in front of it.
Impact The hours the process consumes in a month, and what one error costs when it happens. “Improves efficiency.”
Ease of implementation What already exists: where the documents live, who grants access, what has to be built. “Should be straightforward.”
Technical debt What you will still be maintaining in two years if you build it. Left blank.
Filter 1 — Data owned Where the material lives today and which team controls it. Material you would have to buy, license or extract from a system nobody owns.
Filter 2 — Cost baseline What the process costs today, stated in hours, before any AI touches it. A projected saving with no measured starting point.
Filter 3 — Named acceptor The person who will accept the output as their responsibility. “The business.”

Label the survivors the way Iternal labels them in its own reports: quick wins are easy to implement with the highest return in a short time, strategic bets are harder with a longer path to return. Both belong on the roadmap. Only one of them belongs in your first pilot.

How Many Candidates the Exercise Should Produce

A structured exploration typically returns three to ten candidates, and Iternal is clear about why the range runs wide: the number moves with how much detail goes into the answers. A thin interview returns a thin list. The deliverable is a shortlist by design, built to narrow the field rather than to settle every question about every candidate.

What happens next is the step organizations skip. Leadership takes the ranked grid and picks which candidates go into a deeper look; customers usually narrow it to two, and the two need not be the top two, because a lower-ranked candidate with a willing owner beats a higher-ranked one with none. Then run one. Organizations that start five at once finish none, because the muscle memory is not built yet, and a quick win that lands buys the appetite for something harder.

When Nobody Can Say What Problem This Would Solve

One buyer put it flatly: everybody talks about AI readiness without saying what problem actually gets solved. The objection is fair, it is common, and it usually travels with a sibling — a buyer who does not recognize the need because nobody has ever framed the problem in words they would use. Both are vocabulary failures wearing a technical costume.

The answer to both is a question set rather than a demonstration. Good discovery baselines what success looked like the last time the process ran well, which manual or software tools were involved, and where the gaps sit. Iternal builds prompted question lists into its own tooling for exactly that reason: the useful questions are the ones practitioners forget to ask. Five do most of the work:

  • What did this process look like the last time it went well, and which tools were used?
  • Where does the work stop and wait for a person?
  • Which document does somebody retype into another system?
  • What gets checked a second time because nobody trusts the first pass?
  • Who would notice tomorrow morning if the output were wrong?

Run the set per function, never once per company. A sheet filled in by one person reflects that person's role and misses the operations view, and a guided exercise narrows to a single area — IT ticketing, say — the moment one pain point gets described. Sales, legal and IT each need their own pass, because two organizations in the same industry can carry completely unrelated pain points.

Translate the Example Into the Language of the Team That Owns It

The most common reason a demonstrated example fails to land has nothing to do with the technology. Buyers asked us again and again whether a use case shown to them actually applied to their own situation, and the gap was almost always vocabulary: the example was true, and it was spoken in somebody else's words. One evaluation stalled on exactly that, and the same numbers were worth re-presenting once they were framed in the buyer's language. Trust arrives when the output reproduces the vocabulary a team already reached on its own.

One buyer asked whether AI could help automate the invoice vouchering process for supplier invoices. That sentence is what a generic label sounds like after a department has taken ownership of it. Translating in that direction — from the label to the sentence — takes four lines and about five minutes:

Line The generic version The version the team owns
The name Intelligent document processing Getting supplier invoices vouchered without anyone retyping them
The data Unstructured documents The invoice files suppliers send in, in the formats they send them
Success Higher accuracy Every line on the voucher matches the invoice, or it comes back flagged for a person
The acceptor The business The accounts payable supervisor who signs the batch

The owned version is testable and the generic version is not. Run the same translation on any candidate that arrived from a template. A security team did it for inbound compliance questionnaires and the sentence turned concrete immediately: pull our historic answers into an indexed set and let an agent draft the next questionnaire from them. A candidate that survives translation into a department's own words is one you can scope, price and accept.

What to Do With the Claim That 95% of AI Projects Fail

The number travels well. Its subject does not. Every time the line has appeared in front of us it came from the selling side of the table, never from a buyer describing their own record — and Iternal has run it in its own marketing, so read this as a correction we apply to ourselves first.

Across those tellings the subject drifts: sometimes pilots fail, sometimes projects, implementations, initiatives or investments. A study is named only occasionally. One telling hedged it as “probably 95%”. And no buyer has ever connected the figure, in front of us, to a stalled budget or a canceled program.

Treat it as an unsourced marketing number. Repeat it, if at all, with that scope attached in the same breath, and never let it be the reason you start. The reasons pilots actually stall are specific and recoverable, and buyers described them in their own words — better material for a business case than a percentage nobody can source.

The One-Page Definition of a Pilot Worth Running

A pilot that cannot fail is not a pilot. Four lines on a single page turn an interesting idea into something an organization can accept or reject, and every line has to be filled in before the work starts rather than argued about afterwards:

Line What goes on it The test it has to pass
Data owned The exact material the pilot will run on, and the team that controls it. Somebody can hand it over next week without a new contract.
Cost baseline What the process costs today, in hours. It was measured rather than estimated afterwards. Iternal builds the same comparison into its proof-of-value deliverable, setting today's manual cost against the automated alternative so leadership approves on evidence rather than a promise.
Named acceptor The person who accepts the output as their responsibility. They were asked, and they said yes.
Written success threshold The number that decides it, expressed as accuracy, usability and throughput. It is specific enough to fail. One engagement set zero critical accessibility violations as its proposed criterion — a threshold with a wrong side.

Add one administrative line while the page is open: access and credentials. Write what the delivery team needs, and when, into the statement of work before anything is signed. A pilot that stalls waiting for a login stalls as completely as one that misses its accuracy target, and it stalls more quietly.

What Starting Actually Requires, and How Small You Can Start

Buyers asked the practical version constantly: what do you need from us, and can we dip a toe in first? The answer is shorter than most expect. For a device-side proof Iternal asks for a machine capable of running the models, test licenses, the people who will run it, the devices and chip types they will run it on, and their contact details. Check the machine first. Everything after that is administration.

Small is genuinely available. The standard opening is a staged engagement that builds the first environment and proves it with a subset of users on a subset of the environment — a few hours of your people's time for scoping, then delivery of the results. Collecting the sample documents can run as week zero rather than as a prerequisite that delays everything.

On data, the default and the fallback are different things and both are real. Iternal's proof of value runs on your own documents and shows your own content back to you, because a result on your material is the only result that predicts your result. Where a data-sharing policy blocks that, a mini pilot runs on public or mock material resembling your environment, demonstrating the flow and the end output with nothing confidential in it — the same pattern that lets a program start on public solicitation data and move restricted material in later. Sanitized samples are enough to scope the effort and build the business case. They cannot tell you your own accuracy.

Pin it down: questions for your evaluation
  • Has this exact workload been delivered before, and for which kind of organization?
    Whether you are buying a repeat or a first of its kind, which is the difference between a discount and a premium.
  • Will the pilot run on our own documents, or on representative sample material?
    Whether the result you see predicts your result or only demonstrates the flow.
  • What is the written success threshold, and who signs that it was met?
    The exit criteria, agreed before the work starts rather than negotiated after it.
  • What access and credentials do you need from us, and by when?
    The most common stall, moved into the statement of work where somebody owns it.

Why Pilots Stall Short of Production

Buyers named the problem before we did, and the phrase they used was pilot purgatory. In their own words: a lot of pilots and a lot of POCs frustratingly not scaling into production; AI investment ending in a proof of concept, a demo and impressive slide decks with no shift in the metrics that matter. The stall is not mysterious. A short list of causes accounts for most of it:

  • Criteria written afterwards. Pilots fail for want of the right evaluation criteria, having never been positioned to succeed in the first place.
  • Scope that will only accept perfection. One buyer described an idea that could be pursued only if it arrived as a fully comprehensive enterprise product and architecture on day one.
  • Too many starts. Organizations that begin five at once finish none.
  • Data nobody prepared. If the data is not formatted for AI, the project fails regardless of where the data sits.
  • No route out of the prototype. Companies that ran the prototype successfully then had no scalable production model to move it into.
  • Access nobody owns. A pilot waiting on credentials that no one is funded to chase ends quietly, and usually without a decision.

Each of the three filters closes one of those doors, and the written threshold closes another. Iternal designs around the same failure: the opening engagement is built as the first step into a production arena rather than a paper or conference-room exercise, and its output carries the infrastructure requirements for the production version alongside the result. A pilot designed as a demonstration ends as a demonstration.

Answered elsewhere
FAQ

FAQ: Picking Candidates and Proving the First One

Work backwards from repetitive hours and expensive errors. Any process with a document in, a judgment in the middle and a document out is a candidate, so walk each department against four recurring shapes: enhancing existing content, generating new content, working through long documents, and answering the same quick questions repeatedly. Name each candidate in the words of the team that runs it.

Rank every candidate on impact, ease of implementation and the technical debt it creates, then apply three filters to the top of that list: you already own the data, you can state the current cost in hours, and one named person will accept the output as their responsibility. Anything missing one of the three becomes a demo that never ships.

Translate it before you judge it. Buyers ask that question constantly, and the gap is almost always vocabulary rather than capability — the example was shown in somebody else's words. Rewrite the name, the data, the success test and the acceptor in the language of the team that would own it. A candidate that survives translation is one you can scope and accept.

Yes, small is genuinely available. For a device-side proof Iternal asks for a machine capable of running the models, test licenses, the people who will run it, the devices and chip types involved, and contact details. The standard opening proves the first environment with a subset of users, and collecting the sample documents can run as week zero of the work.

In writing, before the work starts. Iternal groups first-pass criteria into accuracy, usability and throughput; set a number in each, make it specific enough to fail, and name the person who signs that it was met. One engagement used zero critical accessibility violations as its proposed criterion — a threshold with a wrong side.

Your own documents by default: a result on your material is the only result that predicts your result, and Iternal runs its proof of value on customer content for that reason. Where a data-sharing policy blocks it, a mini pilot on public or mock material demonstrates the flow with nothing confidential in it — enough to scope effort, short of proving your accuracy.

Buyers called it pilot purgatory and named the causes: evaluation criteria written after the fact, scope that only accepts a full enterprise architecture on day one, five starts at once so none finish, data never formatted for AI, no production model to move a prototype into, and credentials nobody was funded to chase.

Fill In One Sheet This Week

The whole exercise fits on a page. Name three processes in the words of the people who run them, score each on impact, ease of implementation and technical debt, then run the three filters down the top of the list. Whichever candidate keeps its data, its hours and its named acceptor is your first pilot. The rest are ideas, and ideas keep well.

John Byron Hanby IV
About the Author

John Byron Hanby IV

CEO & Founder, Iternal Technologies

John Byron Hanby IV is the founder and CEO of Iternal Technologies, a leading AI platform and consulting firm. He is the author of The AI Strategy Blueprint and The AI Partner Blueprint, the definitive playbooks for enterprise AI transformation and channel go-to-market. He advises Fortune 500 executives, federal agencies, and the world's largest systems integrators on AI strategy, governance, and deployment.