Skip to main content

Most of your AI agents don’t deserve to exist

AI ROI might seem elusive, but it’s real. It’s just concentrated in a small share of what companies build, and the companies capturing it are the ones willing to kill what doesn’t earn its cost.

That was the throughline of a panel at the 2026 WSJ Tech Council Summit (opens in a new tab), where Retool CEO David Hsu, McKinsey senior partner Kate Smaje, and Glean CEO Arvind Jain joined AI reporter Isabelle Bousquette to talk about enterprise AI spend. David gave one example of a company that had built roughly 500,000 AI agents (more than its own headcount) with no idea where the money was going. Converting the roughly 95% better suited to fixed, deterministic workflows saved an estimated $20 million to $30 million in a year. That’s an extreme case, but the math behind it shows up almost everywhere enterprises are spending on AI.

What enterprise AI ROI actually looks like across the Fortune 500

Across Retool’s customer base, including many Fortune 500 companies, token spend (the underlying unit of AI cost, billed by how much text a model processes) returns something like 3x to 5x on average. But that average is doing a lot of hiding. By David’s estimate, around 90% of tokens spent are net negative. The remaining 10% carries the entire return.

Smaje’s research, gathered independently from a few thousand executives, reached a similar conclusion. More than 80% report real personal productivity gains from AI according to McKinsey (opens in a new tab). It helps them write faster, move through tasks faster, and feel more capable day to day. But when asked whether that shows up as measurable EBIT impact, the number drops to 37%. Ask whether it’s the kind of value that would move investors, and it falls to 6%.

Smaje called the first number “fun with maths”—saving twenty minutes on a task, multiplied across a day, feels significant, but it doesn’t mean a company can cut headcount by 20%, because the time saved was never structured that cleanly to begin with. And so two people, working from two different data sets, arrived at the same conclusion: most of what gets counted as AI value doesn’t survive contact with a P&L.

Why leaders can’t always predict which AI agent will pay off

David offered a comparison from one Fortune 10 retailer that illustrates why the 90/10 split is so hard to see coming. The company built two things around the same time. One was an agent that photographed competitors’ shelf prices in stores and fed the images into a model to inform pricing. The other was a much less impressive-sounding automation that monitored competitors’ earnings calls for new product mentions. The pricing agent—the harder, more expensive build—produced close to nothing. The earnings-call monitor produced an estimated $50 million in value.

Arvind Jain, CEO of Glean, made a similar point that, even when AI is genuinely helping, plenty of departmental use cases never had a clean “before” metric. The ROI of activities like rewritten contract review or faster RFP responses is real, but hard to put a number on. Between that and David’s pricing-versus-earnings-calls split, nobody can reliably guess which use case will pay off by just looking at how sophisticated it sounds. Agent ROI, in practice, gets decided after the build, not before it. When you can’t determine ROI in advance, you have to build enough, watch closely, and be willing to cut what isn’t working.

The cost of poor AI visibility

That “watching closely” part is where most companies stumble. Retool’s State of AI Governance report found just 5% of CIOs, CTOs, and CISOs have full visibility into AI tools in production. One team built an AI-powered out-of-office responder—something email software has handled for free for twenty years—that scanned every message across every Teams channel to decide whether it warranted a reply. It ran up $10,000 a day before anyone noticed, because nobody had a way to see it running until the bill showed up at the end of the month.

But Smaje warned against the instinct to treat rising cost as a reason to shut AI use down, arguing instead for what she calls bending the curve rather than shutting the taps off, so usage keeps climbing while cost per unit of usage comes down. She treats AI as a people problem as much as a technology one, spending as much on retraining and changing how work gets structured as they do on the tools themselves. The technology doesn’t decide to kill the out-of-office bot that’s racking up a hefty bill. The humans in the habit of actually checking what’s running do, and they’ll have to do it enough that a $10,000-a-day mistake doesn’t survive a month.

The discipline that actually drives AI ROI

The retailer couldn’t have known in advance that the earnings-call monitor would outperform the pricing pipeline by that much. Someone had to watch closely enough to tell the difference once both existed, then shut down the loser instead of defending the sunk cost.

If you run technology for a large company, that’s a fair test of where you stand. Pick your five most expensive AI automations right now, and ask whether you could say, specifically, what each one produced last month. If the honest answer is “I know what we spent, not what we got,” then a better model won’t be enough. It’ll be building governance around the model to ensure it’s running well and delivering something real.

Retool lets teams vibe code apps and agents securely, on one platform where IT can see what’s running before the bill arrives. See what you can build in our App Gallery.

FAQs

Published

Category

Insights