Cheaper Per Token Is Not Cheaper Per Task
Opus 5.5 and GPT-6 cut prices 40-50% per token, yet a benchmark task costs the same. Where the tokens go, and how to measure your real cost per job.
Opus 5.5 and GPT-6 cut prices 40-50% per token, yet a benchmark task costs the same. Where the tokens go, and how to measure your real cost per job.
On the morning of September 22, two things happened within hours of each other. Anthropic released Claude Opus 5.5, saying it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. OpenAI released GPT-6 Sol and Luna, priced 50% below GPT-5.6. Latent Space summed the day up as "everybody cuts prices 40-50%", and if you read only the headlines, your AI bill just got cut nearly in half.
It did not. Or rather, it might have, and you will not know until you measure the right thing. Price per token fell. Tokens per task rose. For a lot of real work those two moves cancel out, and for some workloads the second one wins. This post walks through what actually changed, where the extra tokens come from, how to measure your own cost per finished job in an afternoon, and what to change in your setup this week so the price cuts show up on your invoice and not just on the pricing page.
Three labs moved in one week, and it helps to be precise about what each one said.
So two of the three cut list prices by 40 to 50%, and the third held a low price steady. If you buy tokens by the million, the unit price of frontier intelligence dropped by nearly half in a week. That part is real.
The number that matters is not price per token. It is price per finished task. Artificial Analysis ran Opus 5.5 through their Intelligence Index and reported that a task costs $5.98 on Opus 5.5 (max) against $5.86 on Opus 5 (max). Read that twice. The model that is 40% cheaper per token costs slightly more per task on their benchmark, because it spends far more tokens getting to the answer.
That is not a scandal. Newer models often reason longer, retry more and write more, and on hard tasks that extra work is exactly what produces the better result. But it means the pricing page and the invoice are measuring different things. Anthropic priced the token. Artificial Analysis priced the job. You pay for the job.
The same thing shows up in the other direction. Fireworks released Ember-1 this week, a model trained on top of Kimi K3 that delivers the same quality with 40% fewer tokens by learning to cut reasoning that does not change the answer. Their own developers did not notice the switch. No list price moved, and the bill fell 40% anyway. Tokens per task is the lever, and this week both labs pulled it, in opposite directions.
If you are on a subscription plan rather than the API, you never see a token count at all, which makes this harder to notice and easier to be hurt by. Lon measured it after Anthropic made Fable 5 permanently available in subscription plans and the model "felt dumber" to him. Measured five different ways, August delivered dramatically fewer thinking tokens than July. Same model name, same monthly price, less thinking per request.
There is no conspiracy needed to explain that. Effort defaults change between releases, capacity gets managed, and a flat-rate plan gives the provider every incentive to spend fewer tokens on you. The point is simpler: on a subscription, "cheaper" can arrive as more usage before you hit the limit, or as the same usage with less reasoning behind it, and you cannot tell which from the outside unless you count.
So count. Most agent harnesses expose an effort or reasoning setting. Set it explicitly for the tasks that need it instead of trusting the default, and when a model starts feeling worse, check how much it is thinking before you rewrite your prompts.
Benchmarks are someone else’s workload. Your cost per task is a two-hour experiment, and it is the only number worth switching models over. Here is the protocol I would run this week.
You will learn two things. First, whether the price cut is real for your work. Second, which of your tasks are token-hungry, which is where the next section comes in.
Once you can see cost per task, the obvious move is to stop sending every task to the same model. Three patterns are working for small teams right now.
Let the frontier model write the plan and review the result, and let a cheaper model do the typing in between. The Cursor cancellation threads this week were full of people doing exactly this by hand: "I plan with Sol and execute with Luna." Tools like pilotfish wire the same split into Claude Code with a verification step guarding quality. The frontier model’s tokens are the expensive ones, so spend them where judgment matters.
A surprising share of agent spend is classification: is this ticket urgent, which tool should I call next, does this diff break a rule. None of that needs generation. Decision models answer typed questions in one forward pass. Ollaya runs open Jev-style decision models locally; its winnow:e4b scores 0.722 on typed decisions against 0.738 for TypeSafe’s hosted Jev, in 89 ms on an RTX 4090. And NobodyWho showed the mechanic in 25 lines of Python: load any GGUF model, give it lettered choices, read the next-token logits. If you are paying a frontier model to pick between three options, that call is the first one to replace.
Tokens per task include everything the agent reads before it does anything. Fifteen MCP servers with 20 to 30 tools each can burn tens of thousands of tokens before the first word of work. Group servers by job so the agent loads one group at a time, keep instruction files short, and prune the tools you have not called in a month. This is the cheapest optimisation on the list because it applies to every task, every time.
If you pay $20 or $200 a month, the API price cut reaches you indirectly, if at all. The plan’s limits are set in the provider’s cost, so a 40% cheaper model should, in principle, mean more usage before you hit the wall. Whether it does is up to the provider, and this week’s r/cursor threads are a reminder that whoever owns your editor decides which model you get and at what effort.
Two practical moves. First, keep one direct API key even if you live on a subscription, so you can run the five-task experiment above and see real token counts. Second, treat a model that suddenly feels worse as a measurement problem, not a prompt problem. Count thinking tokens first.
This was a good week for anyone who buys intelligence by the token. Frontier models got 40 to 50% cheaper per unit, and that trend is not reversing. But the unit you pay for is the finished task, and the models spending those cheaper tokens are also spending more of them. The builders who win the price war are not the ones who switch to whatever is cheapest on the pricing page. They are the ones who know their cost per job, route each task to the model that earns it, and check the thinking-token count when something feels off.
If you are shipping something and want it in front of builders who care about numbers like these, list it on HelloBuilder. One URL, and the page writes itself.
Get the weekly digest for AI builders & vibe coders. Curated tools, resources, and stories. Skip the scroll.