AI Agenda Live: From Token Maxing to Value maxing: Getting More From Every Unit of Compute
In AI circles, conversations about compute shortages are as common as talk of the weather. As agentic workflows threaten to multiply token consumption further, it’s little wonder that talk has shifted from securing GPUs to optimizing them.
At a recent panel, The Information’s Phoebe Liu sat down with three AI leaders—from chipmakers and AI clouds to Fortune 500s—to discuss how they’re squeezing more tokens from every unit of compute. Those on the panel were:
- Marc Boroditsky, Chief Revenue Officer, Nebius
- Mattie Toia, VP, Infrastructure, Uber
- Dion Harris, Sr. Director, HPC & AI Hyperscale Infrastructure Solutions, NVIDIA
Optimizing to the customer’s KPI
Nebius was recently rated Platinum, the top tier, in SemiAnalysis’ ClusterMAX rating, which puts GPU cloud providers through a suite of hands-on tests on the compute they supply. This distinction, said Chief Revenue Officer Marc Boroditsky, reflects the company’s ability to deliver high performance at high efficiency.
“Not only are we a well-regarded supplier for neolabs, but we’re also enterprise-grade, so ready to supply the entire market,” he said.
He noted that Nebius tailors optimizations in Nebius Token Factory, its inference platform, to specific customer’s specific goals, whether speed, reliability, quality or price.
One recent example of how they do this is speculative decoding, which allows smaller models to predict likely next tokens that a larger model verifies in bulk, rather than generating one token at a time. Nebius uses each customer’s own traffic to improve the result.
NVIDIA, meanwhile, is helping companies avoid unnecessary compute altogether. A feature in its Dynamo platform locates already-computed data sitting in a cluster’s KV cache so teams don’t waste compute recalculating it.
“We’re seeing a lot of bang for the buck in terms of driving performance, efficiency and overall effectiveness,” noted NVIDIA’s Dion Harris.
Uber has seen its token costs stabilize in recent months even as its use of agentic workflows has grown, said Mattie Toia. She credited caching pushed down to individual sub-agents, plus a simpler shift: giving engineers visibility into their own usage.
“It starts by giving visibility to individual users—how many tokens they’re spending, the relative cost of them. That gives people context to think about how they’re optimizing,” she said.
Boroditsky described the shift as moving from “token maxing” to “value maxing”: measuring how many agent turns, tokens and dollars it takes to reach an outcome, whether that’s a document or a completed form.
“What you’re looking for is, how do I actually optimize the value that I’m creating for the cost that I’m spending?” he said.
The ecosystem at work
Improving the economics of compute is a job for the whole industry, not any one company, the panel agreed. “It takes a lot more than a chip to deliver AI at scale,” Harris said.
He pointed to compounding gains from pairing new hardware with better software—kernel tuning, smarter frameworks, and tools built across NVIDIA’s ecosystem. Over a chip’s lifetime, this can drive 3-5X performance gains from day zero to year five.
Boroditsky added that no single model wins every use case.
“Models matter a lot, but by themselves, they don’t solve a problem,” he said. “It’s the entire end-to-end solution.”
Data discipline
One thing companies can do today, Toia said, is get their own data in order.
“There is a lot of valuable data that most organizations have that is really important to grounding your models,” she said. Uber built an internal “context graph”—roughly 24 million nodes pulled from Jira tickets, system design docs and code reviews—to feed its AI agents better context, cutting query times from 20 minutes to 30 seconds while using fewer tokens.
Boroditsky agreed with this approach, noting that the companies that build centralized resources get better results. His five-year prediction: “The leading companies in the market will have taken control of their intelligence.”
The future of pricing
Asked whether Nebius would consider outcome-based pricing, Boroditsky said the company was open to the idea.
“I don’t think we’ve figured out yet the right economic model between the supplier and the customer,” he said.
Nebius already runs capacity auctions and recently launched spot pricing. As the applications built on Nebius move toward outcome-based models, Boroditsky noted that vendors will have to adapt. Looking ahead, he expects managed services to route workloads intelligently on customers’ behalf.
“In the coming quarters, you’re going to see a lot more managed services that are doing more of the heavy lifting,” he said.