As Token Rationing Begins, Think About What Doesn’t Need Tokens
I recently read Every’s newsletter “Token Tightening,” which reports on the shift away from “tokenmaxxing” toward real discernment about where tokens get consumed. Every’s head of tech consulting, Mike Taylor, anticipates that companies will start distributing token budgets (Uber blew through its 2026 token budget in four months), along with access to the highest cost and highest yield models, based on the returns each person can show for their token use. Think of it as a trading desk. Capable engineers get to manage a book many times their salary, with risk limits, auditing, and approvals for the big bets.
The obvious response to all this is to ration tokens. The sharper response is to ask which work never needed tokens in the first place. That second question is the one I want to explore, because I think it reshapes the whole conversation.
Why this is happening
The AI companies eventually have to make money. The leaked, FT-verified financials that Ed Zitron reported tell the story plainly. OpenAI’s operating loss widened from roughly $9 billion in 2024 to nearly $21 billion in 2025. Revenue over that same stretch actually tripled, from $3.7 billion to about $13 billion. So this is not a demand problem. The problem is that the cost of building and serving frontier models is outrunning even that.*
That trajectory is unsustainable, and the underlying cost is what it is for a model-based business. So prices are going up. Token rationing is the first visible symptom. It will not be the last.
The part that the cost squeeze makes obvious
Here is where the new cost consciousness should drive serious interest in a solution like TrueMath. We do not spend tokens to deliver answers. We use a tiny bit of tokenization to convert a natural language prompt into JSON that the TrueMath engine ingests, and we let callers pass structured text or JSON directly, including in the inbound API payload, which bypasses token use altogether.
That is the whole point of a deterministic co-processor for AI workflows. It uses no tokens. It runs on traditional, SaaS-style CPU compute. And it delivers the same answer every time, with an audit trail.
We recently built a cost comparison calculator for a single representative calculation, holding output quality constant and assuming the LLM iterates its way to something like 99.9% accuracy. On that basis, the compute cost of getting an answer from a frontier model like Claude Opus 4.8 runs about 37 times higher than getting it from TrueMath. Speed tells the same story. TrueMath solves in milliseconds, and even after you include the natural-language-to-structured conversion step, it comes back about 19 times faster. Or deliver us structured text or JSON and have an answer before you can blink.
So if token cost is a line item to be controlled and rationed in the modern enterprise, then offloading the functions a CPU performs better, deterministic calculation chief among them, stops being a nice-to-have. It becomes a mandate.
Let’s be honest: LLMs are not TrueMath’s competitors
The use case TrueMath is designed to address is calculations that need to be performed at volume in precisely the same way every time. We’re not built for the occasional math question. So, in reality, we’re not really competing against the notion that “We’ll just do all our math in Claude and it will be fine.” No one we’ve talked to on the client side believes that. No serious person, anyway.
Where we actually compete is against custom builds. Commission a dev team to build out the calculation portion of the app in a way that is deterministic, that converts natural language into structure math prompts, that stores not just the answer, but the method and the underlying assumptions and defaults for future reference by clients, counterparties and auditors. What you’ve done there is take on the burden of building out the math engine, the customization engine, the storage and governance and sharing and collaboration-in-a-trusted-environment engines. And not only building, but paying for dev work to customize every time you need to expand, and paying for the maintenance and bug fixes and edge case investigation when your users hit those edge cases. In other words, you’ve now assumed the cost of building and maintaining TrueMath. This is where the Stripe analogy becomes clear. How many ecommerce sites are rolling their own payment mechanisms today? Why is that?
The CEO of a fast growing AI-based property technology company said it best in an intro meeting I had with them recently. “We’re going to cut three to six months off our dev time getting a [highly calculation-dependent feature that requires 0% error rates] out the door. This is incredible.”
Right. That’s the real comparison. The same CEO said “People just dramatically underestimate the risk of relying on LLMs for repeatable calculations.” Right again.
This discussion is going to speak to a very specific type of product owner. You know who you are. If that’s you, please let me know if you would like to take TrueMath for a spin, see the cost comparison calculator for both TrueMath vs. LLMs and TrueMath vs. custom build, or throw your own math challenges at our engine. We would be happy to give you the demo.
*The headline net loss for 2025 reads closer to $38 billion, though most of that gap is a one-time charge tied to the for-profit conversion, which is why the operating number is the one more relevant to our discussion here.
Reach out: bill.kelly@truemath.ai
Learn more: truemath.ai
Sign up for early access: https://app.truemath.ai/signup
Discover more from TrueMath
Subscribe to get the latest posts sent to your email.