What AI value realization means
AI value realization is the work of converting a model that performs into a business result: the workflow the capability sits in, the decisions it actually changes, the controls around it, and the measure that establishes whether anything improved.
It is a separate question from technical performance, and it has a separate owner. “Does the model do the task?” and “did the business get anything?” are answered by different work. The first is a modelling question. The second is a design question about how people work, who decides, and what is counted.
Most disappointment attributed to AI is disappointment in the second question while the first was the only one anybody was accountable for.
What the evidence shows, and what it does not
Two studies are worth an executive's attention, because they were conducted differently and arrived at the same place.
MIT's Project NANDA reported in July 2025 that 95% of enterprise generative AI pilots produced no measurable return on the profit and loss account, against enterprise investment it put in the range of 30 to 40 billion dollars. The finding rests on a review of more than 300 publicly disclosed initiatives, 52 structured interviews and 153 survey responses gathered between January and June 2025.
RAND Corporation reported in August 2024 that more than 80% of AI projects fail to reach meaningful production — roughly twice the failure rate of IT projects with no AI component. That finding comes from structured interviews with 65 experienced data scientists and engineers.
Both figures deserve their qualifiers, and stating them is not a hedge. The NANDA figure is a review of disclosed initiatives and survey responses, not an audit of financial records. The RAND figure is practitioner testimony rather than a census. Neither is a measured result in the sense this site requires before publishing a case.
What they support is a direction, and they support it independently: two different methods, two different years, two different scopes, one conclusion. That is worth more than either number on its own.
The failures happen before the model runs
RAND traced project failure to five recurring causes.
- Leadership and the technical team did not agree on the problem being solved.
- The data was inadequate for the task.
- The organization pursued the technology rather than a business outcome.
- The infrastructure to deploy and operate a model was not in place.
- The problem was one the technology cannot yet solve.
None of the five is a model-quality problem. Each is a decision taken before any model ran, and four of the five are decisions an executive team makes rather than a technical team.
That is the finding worth carrying out of the research. If the causes precede the model, then improving the model cannot address them, and neither can changing vendor.
Why a working pilot can still produce nothing
A pilot answers one question: can the model perform the task under controlled conditions, with clean inputs and attentive people. That is a real question and worth answering.
It does not answer whether anyone's work actually changes, whether the decision the model informs is subsequently made differently, whether exceptions and failure cases have a defined route, or whether the result can be measured at all. Those are organizational questions.
A pilot designed to prove the model is not constructed to test any of them, so passing it is not evidence that they have been answered. This is the specific reason a convincing demonstration and an unchanged P&L are not a contradiction.
What the work actually is
Name the business consequence first. Establish which decision, operational consequence or financial result makes the change worth making, before specifying a tool. Without it, technical activity becomes the objective, which is RAND's third root cause stated as a project plan.
Redesign the work around the capability. Define the future workflow, the points of human review, what happens to exceptions, who holds which decision, and which roles change. A capability added alongside unchanged work adds workload before it adds value.
Separate business and technical ownership, explicitly. Workflow adoption, controls and outcome measurement have different owners from model development, integration and technical performance. Naming the boundary is what makes the work manageable and the failure diagnosable.
Agree the measure before the build. Define the intended result and the baseline while the work is still being scoped. A measure chosen afterwards is chosen with knowledge of the outcome, which is not measurement.
What separates the deployments that pay
NANDA's account of the minority that did produce value is consistent with the RAND diagnosis, which is part of why both are worth reading.
The deployments that paid were embedded into a workflow that already mattered, rather than offered as a tool beside it. They used systems that retained context and improved with use rather than restarting from nothing at each interaction. They were as often applied to back-office process as to customer-facing work. And they were more often sourced through partnership than built internally.
None of those is a statement about model quality. They are statements about how the work was arranged around the capability, which is the same conclusion arrived at from the failure side.
Where this does not apply
Where the model genuinely does not perform the task well enough, no amount of workflow design substitutes for that, and RAND names it as one of the five causes. The distinction worth testing is whether the capability demonstrably works and fails to produce value, or whether it does not yet work.
The first is a value realization problem. The second is a technical or research problem, and treating one as the other wastes time in both directions.