BOLDGROUPTHAILAND Back to overview
All insights
FIELD NOTE 08 / AI VALUE REALIZATION

Your AI works. The business result has not moved.

The failure is upstream of the model — which is why a better model does not fix it, and why the work starts before anyone chooses one.

01IN BRIEF
  • MIT's Project NANDA reviewed disclosed initiatives and survey responses; RAND interviewed practitioners. Different methods, two different years, and both attribute failure to decisions taken before any model ran.
  • AI value realization is the work between a capability existing and a result arriving: the workflow it sits in, the decisions it changes, the controls around it, and the measure.
  • A pilot proves a model can do a task under controlled conditions. It does not test whether the organization will change the work around it.
  • Name the business consequence first, redesign the work second, and agree the measure before the build rather than after it.
02THE ARGUMENT

There is a version of this conversation that has become uncomfortable in a specific way. The pilot worked. The model does what it was asked to do. The demonstration was convincing. And a year later nobody can point to a line in the accounts that moved.

What AI value realization means

AI value realization is the work of converting a model that performs into a business result: the workflow the capability sits in, the decisions it actually changes, the controls around it, and the measure that establishes whether anything improved.

It is a separate question from technical performance, and it has a separate owner. “Does the model do the task?” and “did the business get anything?” are answered by different work. The first is a modelling question. The second is a design question about how people work, who decides, and what is counted.

Most disappointment attributed to AI is disappointment in the second question while the first was the only one anybody was accountable for.

What the evidence shows, and what it does not

Two studies are worth an executive's attention, because they were conducted differently and arrived at the same place.

MIT's Project NANDA reported in July 2025 that 95% of enterprise generative AI pilots produced no measurable return on the profit and loss account, against enterprise investment it put in the range of 30 to 40 billion dollars. The finding rests on a review of more than 300 publicly disclosed initiatives, 52 structured interviews and 153 survey responses gathered between January and June 2025.

RAND Corporation reported in August 2024 that more than 80% of AI projects fail to reach meaningful production — roughly twice the failure rate of IT projects with no AI component. That finding comes from structured interviews with 65 experienced data scientists and engineers.

Both figures deserve their qualifiers, and stating them is not a hedge. The NANDA figure is a review of disclosed initiatives and survey responses, not an audit of financial records. The RAND figure is practitioner testimony rather than a census. Neither is a measured result in the sense this site requires before publishing a case.

What they support is a direction, and they support it independently: two different methods, two different years, two different scopes, one conclusion. That is worth more than either number on its own.

The failures happen before the model runs

RAND traced project failure to five recurring causes.

  • Leadership and the technical team did not agree on the problem being solved.
  • The data was inadequate for the task.
  • The organization pursued the technology rather than a business outcome.
  • The infrastructure to deploy and operate a model was not in place.
  • The problem was one the technology cannot yet solve.

None of the five is a model-quality problem. Each is a decision taken before any model ran, and four of the five are decisions an executive team makes rather than a technical team.

That is the finding worth carrying out of the research. If the causes precede the model, then improving the model cannot address them, and neither can changing vendor.

Why a working pilot can still produce nothing

A pilot answers one question: can the model perform the task under controlled conditions, with clean inputs and attentive people. That is a real question and worth answering.

It does not answer whether anyone's work actually changes, whether the decision the model informs is subsequently made differently, whether exceptions and failure cases have a defined route, or whether the result can be measured at all. Those are organizational questions.

A pilot designed to prove the model is not constructed to test any of them, so passing it is not evidence that they have been answered. This is the specific reason a convincing demonstration and an unchanged P&L are not a contradiction.

What the work actually is

Name the business consequence first. Establish which decision, operational consequence or financial result makes the change worth making, before specifying a tool. Without it, technical activity becomes the objective, which is RAND's third root cause stated as a project plan.

Redesign the work around the capability. Define the future workflow, the points of human review, what happens to exceptions, who holds which decision, and which roles change. A capability added alongside unchanged work adds workload before it adds value.

Separate business and technical ownership, explicitly. Workflow adoption, controls and outcome measurement have different owners from model development, integration and technical performance. Naming the boundary is what makes the work manageable and the failure diagnosable.

Agree the measure before the build. Define the intended result and the baseline while the work is still being scoped. A measure chosen afterwards is chosen with knowledge of the outcome, which is not measurement.

What separates the deployments that pay

NANDA's account of the minority that did produce value is consistent with the RAND diagnosis, which is part of why both are worth reading.

The deployments that paid were embedded into a workflow that already mattered, rather than offered as a tool beside it. They used systems that retained context and improved with use rather than restarting from nothing at each interaction. They were as often applied to back-office process as to customer-facing work. And they were more often sourced through partnership than built internally.

None of those is a statement about model quality. They are statements about how the work was arranged around the capability, which is the same conclusion arrived at from the failure side.

Where this does not apply

Where the model genuinely does not perform the task well enough, no amount of workflow design substitutes for that, and RAND names it as one of the five causes. The distinction worth testing is whether the capability demonstrably works and fails to produce value, or whether it does not yet work.

The first is a value realization problem. The second is a technical or research problem, and treating one as the other wastes time in both directions.

03SOURCES
  1. The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed

    RAND Corporation — Ryseff, De Bruhl and Newberry · August 2024

    Structured interviews with 65 experienced data scientists and engineers. Practitioner testimony, not a census.

  2. The GenAI Divide: State of AI in Business

    MIT Project NANDA · July 2025

    Review of more than 300 publicly disclosed enterprise initiatives, 52 structured interviews and 153 survey responses, January to June 2025. A review of disclosed initiatives and survey responses, not an audit of financial records. The report has no canonical public address we can link to, and its method has drawn published criticism.

04COMMON QUESTIONS
Does this mean we should stop running pilots?
No. A pilot answers a real question about whether the model can perform the task. The error is treating a passed pilot as evidence that the organizational questions have also been answered, when the pilot was not designed to test them.
How do we tell whether our constraint is the model or the workflow?
Ask whether the capability demonstrably performs the task somewhere, under any conditions. If it does and no result follows, the constraint is in the work around it. If it does not, the constraint is technical and workflow design will not substitute.
Who should own AI value realization?
The business owner of the process being changed, not the technical team. Workflow, decision rights, controls and measurement are theirs to change, and RAND found four of five failure causes sitting in decisions above the technical team.
What should we measure?
The business result the change was meant to produce, defined with its baseline before the work begins. Model accuracy, usage and adoption are useful operational indicators but none of them is a business outcome.
Are those failure statistics reliable?
They are directional rather than definitive. The MIT figure comes from a review of disclosed initiatives and survey responses; the RAND figure from practitioner interviews. Neither is an audit. Their value is that two different methods, in two different years, reached the same conclusion about the cause.
05WHERE THIS LEADS
WHAT SHOULD CHANGE?

Bring the outcome, the constraint, and the accountable decision into the same conversation.

Discuss the outcome that matters