All posts Roi

The Cheapest Token Is the One You Never Spend

PropelFebruary 25, 20266 min read

There is an enormous amount of energy right now going into making tokens cheaper. Cheaper models, open-source alternatives, clever routing of simple tasks to low-cost engines, frontier models reserved for the hard problems. It is a sensible response to a real problem — AI bills have genuinely exploded, and nobody wants to be the company that burned a year's budget in a quarter. But it optimises a number that was never the expensive one.

Two kinds of token spend

When you spend tokens building software, you're really making two different kinds of expenditure, and they could not be more different in what they cost you.

The first kind is the token spent building the right thing. This is pure value. Whatever it costs, you'd pay it again, because it produced exactly what you needed. Making this token cheaper is nice but marginal — you were going to spend it regardless, and it earned its keep.

The second kind is the token spent building the wrong thing. The feature generated flawlessly against a specification that was vague, outdated, or never actually agreed to. The code is clean. The tests pass. And it's worthless, because it isn't what should have been built. This token didn't just cost you its face value. It cost you that, plus the time to discover it was wrong, plus the argument about what right would have been, plus the token you now have to spend building it again. One wrong build is paid for at least twice, and usually more.

Where the wrong build actually comes from

Here is the part the cost-per-token conversation never reaches: in an enterprise, the wrong build is almost never a model failure. It's an agreement failure. Consider a familiar pattern — illustrative, but recognisable. A feature is specified by the business. Architecture quietly assumes a different integration approach. Security's requirement doesn't surface until review. Three parts of the organisation hold three slightly different versions of what's being built, and none of those versions was ever reconciled into one. The agent builds against one of them, faithfully, at full speed. It's wrong against the other two. It's rebuilt. New assumptions surface. It's rebuilt again. The model performed perfectly every single time. The tokens were consumed re-deriving an agreement the organisation never actually captured.

A cheaper model would have produced exactly the same sequence of wrong builds, just billed at a lower rate. That's the thing the routing-and-savings rush misses entirely. You can reduce the unit cost of a wrong build to almost nothing and still have built it three times, because the number of wrong builds was never set by the price of the model. It was set by how badly the organisation disagreed with itself.

The cheapest token is the one you never spend

Not the one you negotiated a better rate on. The one you didn't spend at all, because the wrong build never happened — because the disagreement that would have produced it was caught and resolved before a single token was committed to construction. That token costs zero, and it's the only token spend that's genuinely free.

You don't get there by building cheaper. You get there by agreeing first. The wrong build is not a generation problem to be solved downstream at the model; it's an agreement problem to be solved upstream, among the people and departments who have to converge on what's right. Capture that agreement — make it canonical, signed, and held against the build — before construction begins, and the most expensive category of token spend simply stops occurring. You're not building the wrong thing more cheaply. You've stopped building it.

The number that actually matters

The industry is converging on a metric it's calling intelligence per dollar — the most capability for the least cost. It's a fine metric for choosing a model. It's a useless one for running a delivery, because it has nothing to say about whether the intelligence was pointed at the right target. Cheap intelligence aimed at the wrong build is the worst spend of all, dressed up as efficiency.

The number that matters is cost per correct build. And the largest lever on that number is not the price of a token. It's whether the organisation agreed, properly and provably, on what to build before it built it. That's the lever Propel exists to pull: to make the wrong build expensive to start — caught at the agreement, where it's cheap — rather than expensive to discover, downstream, where it's your annual budget. Lower your cost per token if you can; it won't hurt. But don't mistake it for the saving that matters. The expensive token was never the one you overpaid for. It was the one you should never have spent.

Four numbers, only one of which matters

  • Cost per token. Falling every quarter, and almost entirely outside your control.
  • Tokens per feature. Improving as Cursor, Claude Code and Codex get better at staying on task.
  • Rework rate. The share of AI-generated code rebuilt because the requirement was wrong. Rarely measured.
  • Cost per correct build. The only figure that includes the rework — and agreement, not model pricing, is the largest lever on it.
Now onboarding enterprise teams

Liked this? Come build with us.

Talk to our team about bringing Propel to your organization.