As enterprise adoption of generative AI tools accelerated through late 2025 and into 2026, technology leaders began to hit a frustrating wall. While otheir rganizations poured millions into AI tokens and model subscriptions, corporate leadership and CFOs began pressing for hard evidence that this skyrocketing spend was delivering actual business value. Today, despite widespread integration of developer assistants and automated tools, organizations continue to struggle with determining whether their huge investments in AI tokens translates into meaningful product outcomes, or merely inflated operational costs.
In the initial rush toward AI integration, engineering departments often relied on raw usage metrics—such as token consumption—to evaluate adoption success. However, high token volume quickly proved to be a poor proxy for genuine productivity.
“As we did that, one of the things that we saw is our spend just went through the roof as we adopted that,” said Shams Chauthani, Chief Technology Officer at Tempo.io. “And the question that our CFO started asking us is… ‘What are we getting for all this stuff that we’re doing?'”
Outcome metrics don’t tell the whole story
As the limitations of “token maxxing” became clear, the industry transitioned to tracking output metrics, such as lines of code generated or pull requests submitted, using engineering management tools like Atlassian DX and Jellyfish. While these production metrics gave engineering managers insight into developer activity, they failed to answer executive questions about business value. Generating code faster did not automatically lead to shipping strategic features or improving software quality, and it often penalized developers spending time on critical tasks like resolving technical debt.
“If you’re measuring how many lines of code you wrote, AI is great about writing millions of lines of code very very fast. But ‘Did you actually deliver value or not?’ was the question that was really hard to answer,” Chauthani noted.
Workforce Intelligence platform
To bridge this gap between engineering activity and financial accountability, companies are seeking ways to connect AI spend directly to strategic business units of work. Tempo recently tackled this challenge with the launch earlier this month of its Workforce Intelligence (WFI) platform, to give organizations granular visibility into how AI investments impact product delivery.
Rather than looking at token counts or raw code volume in isolation, WFI correlates token spend data from model providers like OpenAI and Anthropic with GitHub code commits and maps them directly to Jira tickets, epics, and initiatives.
“We basically said, what’s the unit of measure of productivity and product delivery that we’re looking at? And generally, what that is is Jira in our case, or any ticket management system,” explained Chauthani. “If we can tie the dots between what AI spend happened and what ticket was it tied to, we can now all of a sudden get a visibility into [how] this AI spend really drove this outcome for you.”
This level of attribution is becoming essential as AI expenses grow to represent 20% to 30% of overall R&D budgets. According to the Tempo 2026 State of AI report, 91% of technology leaders currently using AI report that they are unable to delegate work to AI and tie it directly to tangible outcomes. By combining AI cost tracking with human labor tracking—a domain Tempo has addressed for two decades—organizations can evaluate which models are most cost-effective for specific tasks, whether refactoring technical debt or building new capabilities.
“Just giving the AI spend is just part of the picture,” Chauthani explained. “You need the human spend and AI spend together, and the ability to roll that information up in a meaningful way, where somebody can actually make decisions off of that.”