The numbers on early enterprise AI are sobering. MIT's 2025 report, The GenAI Divide: State of AI in Business, examined hundreds of enterprise deployments and found that roughly 95% of pilots produced no measurable impact on the P&L. McKinsey's 2025 State of AI survey tells a matching story from the other direction: adoption is nearly universal — 88% of organizations now use AI somewhere in the business — but only about 39% can point to any measurable effect on enterprise-level earnings. Gartner has been predicting the fallout for a while, forecasting that at least 30% of generative AI projects would be abandoned after proof of concept, citing poor data quality, weak risk controls, rising costs, and unclear business value.
Read those three findings together and a pattern shows up. The gap isn't between companies with AI and companies without it. It's between companies that picked a use case with a number attached and companies that picked one because it sounded like the future.
Why the first use case decides the next five
Your first AI project is not really a technology decision. It's a credibility decision. If it lands, the budget conversation for the second one is short. If it produces a well-received demo and nothing else, every subsequent proposal gets weighed against a memory of money spent for a slide deck.
That's why the instinct to start with the most transformative idea in the room is usually wrong. Transformative use cases tend to touch the most systems, the most teams, and the most ambiguous outcomes — which is exactly the profile of a project that takes eighteen months and can't prove what it did. The first one should be chosen for provability, not ambition.
What a use case that pays actually looks like
The candidates worth shortlisting share a recognizable set of traits. When we help clients run this selection, we're screening for:
- A process someone can describe end to end. If nobody can walk you through the current workflow in five minutes, AI won't fix it — it will inherit the confusion.
- Volume and repetition. Value comes from frequency. A task performed 400 times a week produces measurable savings; one performed twice a quarter produces an anecdote.
- A baseline you already measure. Handle time, error rate, days to close, tickets deflected. If you can't state today's number, you will never be able to prove improvement.
- Tolerable error cost. Start where a wrong output is caught and corrected cheaply, not where it triggers a compliance event.
- Data that already exists and is already decent. Gartner names poor data quality as a leading reason projects get abandoned. Do not pick a use case whose prerequisite is a two-year data cleanup.
- An owner with skin in the game. Not an AI task force — the actual manager whose metric moves if it works.
Notice how little of that list is about the technology. The selection criteria that predict success are almost entirely operational.
The best first AI use case is rarely the most exciting one. It's the one where you can already state today's number — and will be able to state a better one in ninety days.
The traps that quietly disqualify a candidate
A few patterns show up over and over in pilots that stall. The "let's build our own" reflex is one: MIT's research found that tools brought in from outside vendors succeeded roughly twice as often as internally built ones, largely because internal builds absorb the learning curve and the maintenance burden at the same time. The enterprise-wide chatbot is another — broad, unowned, and impossible to attribute value to, which is why so many of them end up as expensive intranet search.
Then there's the subtler trap: choosing a use case that saves time without a plan for what the time gets used for. Twenty minutes returned to forty people is meaningless unless it's redeployed into work that shows up somewhere. Savings that aren't captured are savings that don't exist on the P&L, and that's precisely the disconnect MIT's numbers are measuring.
Run the selection like a portfolio decision
Practically, this is a short exercise, not a strategy engagement. Get eight to twelve candidate use cases on the table from the people who do the work, not just the leadership team. Score each on two axes: how measurable the value is, and how hard it is to implement given today's data and today's systems. Pick from the high-measurability, low-difficulty quadrant even if it feels unglamorous. Write down the baseline metric before you start, name the person accountable for it, and set a review date roughly ninety days after go-live.
Then treat it like any other change: the model is the easy part, and adoption is where the value actually gets realized or lost. A well-chosen use case with no one reinforcing the new way of working produces the same result as a badly chosen one — it just gets there more politely.
The organizations closing the gap aren't the ones with better models. They're the ones that were disciplined about the first question: what specific number are we trying to move, and how will we know we moved it?
