hellobuilder

Command Palette

Search for a command to run...

← Back to blog

Cheaper Per Token Is Not Cheaper Per Task

Opus 5.5 and GPT-6 cut prices 40-50% per token, yet a benchmark task costs the same. Where the tokens go, and how to measure your real cost per job.

Nishant Modi
September 28, 2026 · 8 min read

On the morning of September 22, two things happened within hours of each other. Anthropic released Claude Opus 5.5, saying it performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. OpenAI released GPT-6 Sol and Luna, priced 50% below GPT-5.6. Latent Space summed the day up as "everybody cuts prices 40-50%", and if you read only the headlines, your AI bill just got cut nearly in half.

It did not. Or rather, it might have, and you will not know until you measure the right thing. Price per token fell. Tokens per task rose. For a lot of real work those two moves cancel out, and for some workloads the second one wins. This post walks through what actually changed, where the extra tokens come from, how to measure your own cost per finished job in an afternoon, and what to change in your setup this week so the price cuts show up on your invoice and not just on the pricing page.

What changed on the price sheet

Three labs moved in one week, and it helps to be precise about what each one said.

  • Anthropic: Opus 5.5 is "the first model in our new Claude 5.5 family". It performs at Fable 5.1 level on most tasks and costs 40% less to run than Opus 5. Anthropic credits efficiency work for the cut, and Theo Browne observed on X that the model appears to be smaller than Opus 5.
  • OpenAI: GPT-6 Sol and GPT-6 Luna bring much of GPT-6 Astra’s strengths into faster, cheaper models, launched at 50% below GPT-5.6. Product Hunt’s listing put it as "frontier AI intelligence, now at half the price".
  • xAI: Grok 4.7 shipped the day before at the same price as Grok 4.6, $2 per million input tokens and $6 per million output, and xAI pitched it as "half the price of comparable models" rather than as a cut.

So two of the three cut list prices by 40 to 50%, and the third held a low price steady. If you buy tokens by the million, the unit price of frontier intelligence dropped by nearly half in a week. That part is real.

The catch: tokens per task went up

The number that matters is not price per token. It is price per finished task. Artificial Analysis ran Opus 5.5 through their Intelligence Index and reported that a task costs $5.98 on Opus 5.5 (max) against $5.86 on Opus 5 (max). Read that twice. The model that is 40% cheaper per token costs slightly more per task on their benchmark, because it spends far more tokens getting to the answer.

That is not a scandal. Newer models often reason longer, retry more and write more, and on hard tasks that extra work is exactly what produces the better result. But it means the pricing page and the invoice are measuring different things. Anthropic priced the token. Artificial Analysis priced the job. You pay for the job.

The same thing shows up in the other direction. Fireworks released Ember-1 this week, a model trained on top of Kimi K3 that delivers the same quality with 40% fewer tokens by learning to cut reasoning that does not change the answer. Their own developers did not notice the switch. No list price moved, and the bill fell 40% anyway. Tokens per task is the lever, and this week both labs pulled it, in opposite directions.

Thinking tokens are the line item nobody looks at

If you are on a subscription plan rather than the API, you never see a token count at all, which makes this harder to notice and easier to be hurt by. Lon measured it after Anthropic made Fable 5 permanently available in subscription plans and the model "felt dumber" to him. Measured five different ways, August delivered dramatically fewer thinking tokens than July. Same model name, same monthly price, less thinking per request.

There is no conspiracy needed to explain that. Effort defaults change between releases, capacity gets managed, and a flat-rate plan gives the provider every incentive to spend fewer tokens on you. The point is simpler: on a subscription, "cheaper" can arrive as more usage before you hit the limit, or as the same usage with less reasoning behind it, and you cannot tell which from the outside unless you count.

So count. Most agent harnesses expose an effort or reasoning setting. Set it explicitly for the tasks that need it instead of trusting the default, and when a model starts feeling worse, check how much it is thinking before you rewrite your prompts.

Measure your own cost per job in an afternoon

Benchmarks are someone else’s workload. Your cost per task is a two-hour experiment, and it is the only number worth switching models over. Here is the protocol I would run this week.

  1. Pick five real tasks from your last week, not toy prompts. A refactor you actually shipped, a bug you actually fixed, a doc you actually wrote. Save the exact starting state (a git branch per task works).
  2. Run each task on the model you use today and on the new one, from the same starting state, with the same instructions. Do not tune the prompt for the new model; you are measuring the swap, not your prompt skills.
  3. Log input tokens, output tokens and thinking tokens per run. The API returns them. For Claude Code, Codex and Gemini CLI, tools like tokentab read the session logs and break spend down by model and project.
  4. Record whether the task actually finished, and how many retries or follow-up prompts it took. A run that needed three corrections is one task, not one run.
  5. Compute cost per completed task, not cost per run and not cost per token. Then compare. If the new model is 40% cheaper per token and 60% more verbose, you already know how that ends.

You will learn two things. First, whether the price cut is real for your work. Second, which of your tasks are token-hungry, which is where the next section comes in.

Route by task, not by loyalty

Once you can see cost per task, the obvious move is to stop sending every task to the same model. Three patterns are working for small teams right now.

Plan with the expensive model, execute with the cheap one

Let the frontier model write the plan and review the result, and let a cheaper model do the typing in between. The Cursor cancellation threads this week were full of people doing exactly this by hand: "I plan with Sol and execute with Luna." Tools like pilotfish wire the same split into Claude Code with a verification step guarding quality. The frontier model’s tokens are the expensive ones, so spend them where judgment matters.

Stop paying an LLM to make a decision

A surprising share of agent spend is classification: is this ticket urgent, which tool should I call next, does this diff break a rule. None of that needs generation. Decision models answer typed questions in one forward pass. Ollaya runs open Jev-style decision models locally; its winnow:e4b scores 0.722 on typed decisions against 0.738 for TypeSafe’s hosted Jev, in 89 ms on an RTX 4090. And NobodyWho showed the mechanic in 25 lines of Python: load any GGUF model, give it lettered choices, read the next-token logits. If you are paying a frontier model to pick between three options, that call is the first one to replace.

Cap the context your agent starts with

Tokens per task include everything the agent reads before it does anything. Fifteen MCP servers with 20 to 30 tools each can burn tens of thousands of tokens before the first word of work. Group servers by job so the agent loads one group at a time, keep instruction files short, and prune the tools you have not called in a month. This is the cheapest optimisation on the list because it applies to every task, every time.

What "cheaper" means on a subscription

If you pay $20 or $200 a month, the API price cut reaches you indirectly, if at all. The plan’s limits are set in the provider’s cost, so a 40% cheaper model should, in principle, mean more usage before you hit the wall. Whether it does is up to the provider, and this week’s r/cursor threads are a reminder that whoever owns your editor decides which model you get and at what effort.

Two practical moves. First, keep one direct API key even if you live on a subscription, so you can run the five-task experiment above and see real token counts. Second, treat a model that suddenly feels worse as a measurement problem, not a prompt problem. Count thinking tokens first.

What to change this week

  • Run the five-task experiment on your current model and on Opus 5.5 or GPT-6 Sol. Decide on cost per completed task, nothing else.
  • Set effort explicitly for the tasks that need reasoning, and stop trusting the default.
  • Find the classification calls in your agents and move them to a decision model, local or hosted.
  • Split planning from execution: frontier model for the plan and the review, cheaper model for the typing.
  • Cut startup context: group your MCP servers, shorten instruction files, drop unused tools.
  • Keep an API key alongside your subscription so you can always see the token counts.

The honest read

This was a good week for anyone who buys intelligence by the token. Frontier models got 40 to 50% cheaper per unit, and that trend is not reversing. But the unit you pay for is the finished task, and the models spending those cheaper tokens are also spending more of them. The builders who win the price war are not the ones who switch to whatever is cheapest on the pricing page. They are the ones who know their cost per job, route each task to the model that earns it, and check the thinking-token count when something feels off.

If you are shipping something and want it in front of builders who care about numbers like these, list it on HelloBuilder. One URL, and the page writes itself.

AI is moving fast. Don't get left behind.

Get the weekly digest for AI builders & vibe coders. Curated tools, resources, and stories. Skip the scroll.

Keep reading