By: Natalie Johnson
Boards approved AI budgets on faith for three years, and the bill is now coming due. Finance chiefs who once accepted a pilot deck and a productivity projection are asking a harder question: where, precisely, did the money show up in the numbers? Most organizations cannot answer it, because they never built the measurement chain that would let them.
David D. Ellison, who spent eight years building an enterprise AI business at a global technology company from its earliest stages, argues the failure happens long before deployment. It happens at the evaluation stage, when companies assess the model instead of the system required to turn that model into economic value. His framework is uncomfortable for anyone hoping AI will justify itself after the fact, and that is the point.
Start With The Economics, Not The Model
Ellison’s evaluation runs across five dimensions:
- Economic value
- Technical and data feasibility
- Workflow impact
- Adoption
- Total cost to scale
The financial case has to carry everything, not just development. Integration, infrastructure, security, governance, monitoring, change management, and ongoing operations all belong in the denominator. He points to the National Institute of Standards and Technology’s (NIST) AI Risk Management Framework, which reinforces the need to evaluate AI within the broader organizational, operational, governance, and risk context rather than focusing only on model performance. That framing matters because most internal business cases quietly treat the hard costs as someone else’s line item.
His value equation is deliberately conservative: opportunity size multiplied by achievable improvement, multiplied by adoption, multiplied by probability of success, minus the full lifecycle cost. Three of those four terms are discounts. “That prevents us from confusing theoretical value with value we can realistically capture,” Ellison says. He also insists on a counterfactual before anything scales. What would have happened without the AI? He cites Netflix’s recommender-system work as a model for tying algorithmic improvements to controlled experiments and measurable business outcomes, rather than offline model metrics alone. Without that baseline, a company cannot distinguish correlation from value creation, and every favorable quarter becomes evidence for whatever it wants to believe. As Ellison puts it, “The central question is not, ‘Can we build it?’ It is, ‘If we build it and people use it, what economically changes?'”
The Pilot Extrapolation Trap
The most common failure Ellison sees is structural. An organization identifies a large theoretical opportunity, runs a pilot with impressive model performance, and extrapolates those results across the enterprise. Value does not scale that way. Data integration, workflow redesign, security, governance, employee behavior, incentives, operating costs, and organizational change all determine how much of the theoretical opportunity survives contact with the business. “I learned that the model is often the easiest part,” he says. “The difficult part is changing how work gets done around it.” Economic research on general-purpose technologies supports him, showing that new technologies typically require complementary investments in processes, skills, and systems before their productivity potential is realized.
Ellison is especially skeptical of the arithmetic that multiplies pilot productivity gains by headcount. Research on generative AI in customer-support work found meaningful productivity improvements, but the effects varied substantially by worker experience and skill level. Average gains in a controlled pilot say little about what a 10,000-person organization will capture. His screen for separating a fundable use case from a promising-looking one runs four tests.
- Is there a meaningful economic problem, quantified in what it costs today?
- What decision or action changes because of the AI? A prediction that doesn’t change workflow has little economic value.
- Can the impact be proven through an A/B test, controlled rollout, matched population, or comparable counterfactual?
- Can it scale economically? A system that works for 500 users may collapse at 50,000 once compute, integration, support, and governance are counted.
Klarna’s public reporting on its AI customer-service deployment is, in Ellison’s view, an example of evaluating AI at operational scale through workload handled, resolution time, and cost impact.
Where Did The Capacity Go?
The pressure from chief financial officers to tie AI to revenue or margin rather than productivity is, Ellison believes, correct. Productivity is an operational benefit. It is not automatically financial return. “If AI saves 100,000 hours, I do not automatically multiply those hours by an employee’s salary and declare savings,” he says. The question is what the organization did with the recovered capacity. Avoided a hire. Cut overtime. Served more customers. Increased sales. Accelerated a launch. Reduced cost-to-serve. Absent one of those, the hours are real and the return is theoretical. His three-level measurement chain forces the connection: technical metrics feed operational metrics, which must land on revenue, margin, cost, working capital, or another financial outcome. McKinsey’s research on enterprise AI adoption shows many organizations reporting operational or productivity benefits, while a much smaller group can demonstrate material impact on enterprise financial performance.
What separates the two groups over the next two years will be whether AI is managed as a portfolio of experiments, or as a repeatable capability for changing the business. Ellison expects the winners to redesign workflows rather than insert AI into existing ones, a practice McKinsey associates with stronger reported financial impact. They will set financial measures before development begins, build reusable data and AI platforms, invest in adoption and workforce change, and kill projects whose economics fail fast. They will also be explicit about which kind of value each initiative is meant to produce, whether revenue, margin, risk reduction, customer experience, or strategic capability, and how it will eventually be measured. Accenture’s research similarly emphasizes that broader enterprise value from generative AI depends on data foundations, leadership alignment, operating-model change, and workforce reinvention. Ellison’s conclusion cuts against most current AI strategy. “Models will become increasingly accessible,” he says. “The competitive advantage will be the organizational ability to repeatedly turn AI into changed decisions, changed workflows, adoption, and measurable economics.”
Follow David D. Ellison on LinkedIn for more insights on AI investment evaluation, value measurement, and scaling AI into measurable business outcomes.





