hellobuilder

Command Palette

Search for a command to run...

← Back to blog

Your Coding Agent Needs a Hard Budget Cap

Coding agents make it easy to ship things that cost money. How to put a hard ceiling on every layer, from cloud bills to the agent loop itself.

Nishant Modi
October 6, 2026 · 7 min read
Featured image: Your Coding Agent Needs a Hard Budget Cap

On 3 October, Simon Willison published a short post with a long title: We're going to need default hard budget caps on pretty much everything. His argument is simple. Pay-per-use services should stop when they reach a limit you set, not send you an email about it, and that stop should be the default rather than a setting you have to go and find.

Why it matters now is in one line of his post: "Coding agents, and personal agents (coding agents wrapped in a less threatening UI), greatly reduce the friction of spinning up code that can do useful things." Useful things cost money. A database, a serverless function, a model API call inside a loop: each one is pay-per-use, and an agent can create all three before lunch.

This post is the practical follow-up. Below: why agents change the math on spend, what a cap looked like when one fired on us, the five layers where a cap belongs, and a 30-minute audit you can run on your own stack today.

An alert is an email. A cap is a circuit breaker.

Most billing safety on the internet is built around alerts. You set a budget, the provider emails you at 50%, 80% and 100%, and the meter keeps running. That model assumed a human was in the loop for every change, and that spend grew slowly enough for a person to notice.

Simon describes the failure mode exactly: nobody wants to wake up to a midnight email about a budget limit and find that a rogue service has already consumed "several hundred (or several thousand) more dollars". By the time you read the alert, the money is gone.

Agents break both assumptions. They create infrastructure on your behalf, they write loops that call paid APIs, and they retry when something fails. A retry loop at 3am does not get tired, and it does not read email. A cap is different in kind: when the limit is reached, the service stops. You lose availability instead of money, and you get to choose that trade in advance.

The big clouds are starting to move. According to Simon’s post, AWS launched a monthly spend limit for projects on 16 September 2026 that pauses a project for the rest of the month once it reaches the limit, and Google Cloud launched Spend Caps in July, which put a monthly financial cap on specific services in a project. Neither is the default, which is his whole point.

What a cap looks like when it actually fires

We hit one ourselves. On 25 September, Supabase emailed to say the HelloBuilder project had used 8.21 GB of bandwidth against a 5.5 GB allowance on its plan, and that requests to the project would be dropped until the quota refilled on 2 October.

That is a hard cap doing exactly its job. There was no surprise bill. There was also a database that stopped answering, for a product that at the time had one signed-up user. The cap did not explain why we had used the bandwidth. It just stopped the bleeding and gave us a reason to look.

The cause was ordinary. List pages pulled full content bodies they never displayed, CMS pages rebuilt more often than they needed to, and images were not cached for long. The fix was three changes: lists stopped pulling bodies, CMS pages rebuild hourly, and images are cached for a month. None of it needed more money. It needed attention, and the cap is what forced it.

That is the honest trade-off of a hard cap: it converts a billing problem into an availability problem. For a side project or an early product, that is almost always the trade you want. For a production system with paying customers, you want the cap high enough that it only fires on something genuinely broken, plus alerts well before it.

The cap stack: five places a limit belongs

One cap at the cloud account is not enough, because expensive mistakes rarely start there. They start in a loop, a key or an endpoint. Think of it as a stack, from the outside in.

1. The account: a hard stop on every pay-per-use provider

Every service with your card on file should have either a hard spend limit or a written reason why not. Check whether your provider offers a true stop or only an alert, because the two often sit under the same word in the dashboard. Where a hard stop exists, turn it on. Where it does not, that service is the first place to add the layers below.

2. The key: one API key per project and per agent

Model API spend is where agents burn money fastest, and a single shared key makes a runaway impossible to contain. Give each project, and ideally each long-running agent, its own key with its own limit. OpenRouter lets you put a credit limit on an individual key, and Anthropic’s Console lets you set spend limits per workspace. When one agent misbehaves, it exhausts its own allowance and nothing else.

3. The loop: a maximum on turns and spend per task

Every agent loop you run should carry two numbers that are checked before each model call: a maximum number of iterations and a maximum spend for the task. If you run Claude Code headless, the CLI reference documents both: --max-turns limits the agentic turns of a print-mode run, and --max-budget-usd sets the most it may spend on API calls before stopping. In your own agent code, keep a running total of tokens per task and stop when it crosses the budget, the same way you would stop a recursive function that has no base case.

Measure it as well. Tools like agentacct break each agent task into steps with tokens and estimated cost, which is how you find the one task that quietly costs ten times the others.

4. The endpoint: per-user quotas on anything that calls a paid API

If your product calls a model on behalf of users, every one of those endpoints is a meter that someone else controls. Add a daily quota per user and a stricter one per anonymous visitor, and rate-limit signups if a free account unlocks paid calls. A free tier that wraps an expensive model is one of the first things a scraper finds.

5. The alert: still useful, as the last layer

Alerts are not useless. They are just not a cap. Keep them, set them at 50% and 80% of each cap so you hear about a trend before the stop, and send them somewhere you actually look. An alert in an inbox you check weekly is decoration.

Choosing between a bill and downtime

Every cap forces a decision: when this limit is reached, would you rather pay or go down? Make that decision per service, on purpose, instead of inheriting whatever the provider chose for you.

  • Side projects and experiments: cap low and accept downtime. The worst case should be a paused project, never a four-figure bill.
  • Early products with a few users: cap at two to three times normal monthly spend, with alerts at 50% and 80%. A cap that fires is a bug report.
  • Revenue-critical production: cap high enough that only a runaway reaches it, and invest in the loop and endpoint layers so the runaway never starts.
  • Model APIs used by agents: always cap per key, whatever the stage. This is where spend moves fastest.

Teach your agent the rule

Simon ends with a suggestion that is easy to act on: agents themselves should recommend providers with hard caps and warn when they are about to wire in one without. You do not have to wait for agent vendors to build that. Put it in your AGENTS.md or CLAUDE.md:

"Before adding any paid service, tell me how it is priced and whether it supports a hard spend cap. Prefer services that do. Never add an uncapped pay-per-use service, or a new API key without a limit, without asking me first."

Two sentences in a file the agent reads every session. It will not catch everything, but it moves the decision to the moment it is cheapest to make: before the service exists.

A 30-minute audit you can run today

  1. List every service with a card on file. Your card statement is a better source than your memory.
  2. For each one, write down whether it has a hard cap, an alert, or nothing at all.
  3. Turn on every hard cap that exists, at the level your stage calls for.
  4. Split shared model API keys into one key per project, each with its own limit.
  5. Add a max-turns and a max-spend check to every agent loop you own.
  6. Put per-user quotas on every endpoint that calls a paid API.
  7. Lower one cap until it fires, on purpose, so you know what your product does when it happens.

The last step matters most. A cap you have never seen fire is a cap whose consequences you have not met yet.

Practical takeaways

  • Alerts tell you after the money is gone; caps stop the spend. Use both, and never confuse them.
  • Agents remove the friction that used to slow spending down, so the limits have to live in the system, not in your attention.
  • Cap at five layers: account, key, loop, endpoint and alert.
  • A cap trades a bill for downtime. Choose that trade per service, deliberately.
  • Write the rule into the file your agent reads, so it asks before adding anything uncapped.

Conclusion

Hard caps are not exciting, and that is the point. The best result of this post is a quiet month in which nothing happens. The big clouds are starting to ship real spend limits, but the defaults still favour the meter. Until they change, the caps are your job, and an afternoon spent on them costs far less than the bill that would otherwise teach the lesson.

For more practical notes like this from people building with AI, get the HelloBuilder newsletter. It lands every Monday.

AI is moving fast. Don't get left behind.

Get the weekly digest for AI builders & vibe coders. Curated tools, resources, and stories. Skip the scroll.

Keep reading