IBM announced Apptio AI Value and ROI on August 6th. The product connects AI token spend, labor investment, and infrastructure costs to measurable business outcomes across revenue, cost, speed, productivity, and risk. One client using the platform cut AI spending by 50% without reducing output, freeing capital to fund additional initiatives.
IBM did not build this tool because the market was too sophisticated. It built this tool because the market has a measurement problem.
Gartner estimates that 84% of finance leaders have not been able to measure the ROI of their AI initiatives. Among those who tried, roughly two in five succeeded in quantifying returns. That means about 84% of enterprise AI investment is operating without a financial feedback loop.
This is not an acceptable position in 2026.
The investment numbers are not small. Battery Ventures’ enterprise technology spending survey found that 98% of respondents are increasing their AI spend over the next 12 months. Not a single CXO surveyed is cutting back. Meanwhile, MIT’s NANDA research found that only 5% of enterprise generative AI pilots reach measurable P&L impact. The other 95% stall.
Those two facts together describe an industry spending more money, faster, with less measurement than any comparable technology cycle in recent history.
This is the AI accountability gap. It is distinct from the AI productivity gap, though the two are related. The productivity gap is about whether AI is generating value. The accountability gap is about whether anyone knows.
Finance teams are starting to close the gap from their side. The ECI Research 2026 Application Development survey found that 58.2% of organizations will increase AI governance spending by 10% to 25% this year. That’s not governance theater, it’s finance departments responding to a real problem: they are being asked to approve AI budgets they cannot audit.
IBM’s Apptio product is one commercial response to this. The approach is worth examining regardless of vendor preference because it reflects how the measurement problem needs to be structured.
The core architecture is: token costs connect to initiative costs, initiative costs connect to business outcome metrics, business outcome metrics connect to baseline and target tracking over time. That is a standard capital investment framework applied to AI deployment. The novelty is that AI spending, specifically token consumption, has been opaque enough that this connection could not be made cleanly before.
Token costs are real operating costs. A large language model interaction is not free. At scale, token spend becomes a line item that needs the same financial discipline as compute, storage, or headcount. Organizations that treat AI as a software license expense are going to be caught off guard when they audit where the actual spend is flowing.
The second component is the initiative-to-outcome connection. This is where most enterprises are currently failing. They know what AI tools they are paying for. They do not know which business outcomes those tools are moving. Without that connection, there is no basis for prioritization, no basis for reallocation, and no feedback signal for improving AI deployment decisions.
IBM’s product requires organizations to define proof metrics upfront: cycle time, cost avoided, conversion rate, incident volume. These are not novel metrics. They are the same metrics finance teams use to evaluate any capital project. The discipline of defining them before deployment, not after, is what creates the accountability loop.
For enterprise technology and finance leaders, there are three practical implications.
First, every AI initiative should have a defined proof metric before funding approval. If the initiative cannot be connected to a measurable business outcome in advance, that is a signal the initiative is not ready to fund, not a signal to fund it and measure later.
Second, token spend needs to be in the financial operating model. It is not an IT line item to be managed at the department level. It is an operating cost that scales with usage and requires the same financial oversight as any other variable cost.
Third, the 84% who cannot measure AI ROI are not failing because measurement is impossible. They are failing because measurement was not built into the deployment architecture. This is a solvable problem, but it requires treating AI deployment as a financial management challenge from the start, not an IT challenge with a finance review at the end.
The Gartner finding that 84% of finance leaders cannot measure AI ROI is not a data point about AI capability. It’s a data point about organizational discipline.
The accountability gap will close. The question is whether it closes because organizations build measurement discipline into their AI programs, or because a budget crisis forces the question.
One of those is a lot cheaper than the other.
Robin Green is an enterprise technology executive and the author of The Intelligence Loop. He is Chief Revenue Officer at Occams Advisory.
Get new articles from Robin Green delivered directly.
Insights on AI leadership, the future of work, and human collaboration. No noise. Unsubscribe any time.
Subscribe free