# What are spend caps for AI agents, and how do they work? (/blog/what-are-spend-caps-for-ai-agents)

![Illustration: A small friendly robot with a round camera-eye head drops coins into a plain unlabeled glass jar.](/blog/what-are-spend-caps-for-ai-agents/hero.webp)

A spend cap is a limit on how many credits an agent may spend before the platform refuses further metered calls. CoreSpeed has two. The organization wallet has a billing threshold, which answers `payment_required`. An API key can carry a monthly cap, which answers `key_spend_limit_exceeded`. Both refuse a call before the tool runs.

**Caps are checked before the tool runs**

Every metered call takes one path. The hold check is where the wallet threshold and the key cap apply.

**Authenticate** (member or agent) → **Resolve tools** (what this caller sees) → **Check holds** (wallet threshold, key cap) → **Execute** (the tool runs) → **Charge and record** (one ledger, one trail)

A refused call answers 200 with isError true and the code payment\_required or key\_spend\_limit\_exceeded. Discovery, account reads, key management and memory keep working.

## Why do agents need a spend cap? \[#why-do-agents-need-a-spend-cap]

An agent decides its next call from the last result. A loop that keeps searching, keeps scraping, or keeps regenerating an image has no natural stop. A person at the keyboard would notice the bill. An unattended job you run elsewhere would not. The cap is the stop that does not depend on anyone watching.

On CoreSpeed every metered action is priced in credits, and the per-action prices are on the [pricing page](/pricing). Two examples: `web__search` costs 14 credits per request and covers up to 10 results, and `web__scrape` costs 5 credits per page. Memory operations and discovery (`tools/list`) cost 0 credits. Every charge lands itemized in the organization ledger, so a cap is a number you can check against real history.

## How does the organization boundary work? \[#how-does-the-organization-boundary-work]

Every metered call takes one path: authenticate, resolve the tools this caller can see, check holds, execute, charge, record activity. The hold check is where the wallet threshold applies. When the wallet has crossed its billing threshold, later metered calls are refused before the underlying tool executes.

On `/mcp` that refusal is a tool-level result, HTTP `200` with `isError: true` and the code `payment_required`. The tool did not run. Discovery, account reads and key management keep working, so an agent can still list tools and a member can still manage keys.

Adding credit clears the state. Credit enters the wallet three ways: 3,000 free credits when your first organization is created, no card required, expiring 90 days after the grant; 10,000 credits each month on CoreSpeed Pro, reset monthly; and top-ups of $5 to $1,000 on Pro, which never expire. The Free plan has no top-up.

One state looks similar and is different. `org_suspended` is administrative. Adding balance does not clear it.

## How does a per-key monthly cap work? \[#how-does-a-per-key-monthly-cap-work]

A monthly spend cap can be attached to an API key: in Dashboard → API keys, with the `manage__keys_*` tools from a signed-in agent, or with `cs keys create`. Once the month's recorded usage on that key reaches the cap, later metered calls using the key are refused with `key_spend_limit_exceeded`. The refusal lasts until the month rolls over or someone raises the cap.

Caps attach to keys. The browser sign-in path has only the wallet stop. This is why an unattended agent you run elsewhere should hold its own agent key (`sk-csa-...`): the key has a cap of its own, the agent has an identity of its own, and the activity trail shows what it spent. Keys are per environment, so each environment gets its own key and its own cap.

| Boundary         | Attaches to      | Code on refusal            | What clears it                          |
| ---------------- | ---------------- | -------------------------- | --------------------------------------- |
| Wallet threshold | The organization | `payment_required`         | Plan credits or a top-up                |
| Monthly cap      | One API key      | `key_spend_limit_exceeded` | Month rollover or a raised cap          |
| Suspension       | The organization | `org_suspended`            | An administrative action, never balance |

## What does the agent see? \[#what-does-the-agent-see]

Both refusals arrive on `/mcp` as HTTP `200` with `isError: true`. The code is in `structuredContent.error.code`:

```json
{
  "result": {
    "isError": true,
    "structuredContent": { "error": { "code": "key_spend_limit_exceeded" } }
  }
}
```

Treating a `200` as success is the most common integration bug. Branch on `isError` first. Neither code clears on retry; retrying only adds noise to the activity trail. The right move is to report the code and stop. A member fixes the wallet at Dashboard → Billing, and the key's owner or an org admin raises the cap. The [errors page](/docs/errors) describes both error layers.

## Why is a cap a boundary rather than a reservation? \[#why-is-a-cap-a-boundary-rather-than-a-reservation]

Spend caps are request-boundary guardrails. The check happens when a call arrives. An action already in flight can finish above the remaining amount before its final cost is recorded. A media generation that started below the line completes and is charged at the model's full rate. The next metered call is the one refused.

This matters for planning. Treat a cap as the number past which the next call fails, and keep a margin for the one call that may already be running. For a job that will make many billable calls, start with a small batch and read the cost in Dashboard → Activity before scaling up.

## What keeps working during a hold? \[#what-keeps-working-during-a-hold]

A hold is narrow by design. Discovery still answers, so `tools/list` is unchanged. Account reads still answer. Key management still works, which is how the fix gets applied. Memory operations are free, so an agent can still search and save memory. Only metered calls are refused.

The bill and the trail stay in step. The billing ledger is the authority for money; activity records reference a charge when one exists, and the dashboard joins the two. An audit retry cannot duplicate a charge. So the cost you read in Activity is the cost in the ledger, and the number a cap is measured against. The [billing page](/docs/billing) and the [activity page](/docs/activity) describe each side.

## FAQ \[#faq]

**Does a cap stop reads?** A cap applies to metered calls, whatever they do. A metered read such as `web__search` is refused; free operations such as memory and `tools/list` continue.

**Can a browser sign-in have a monthly cap?** No. Caps attach to API keys. A member signed in through the browser is bounded only by the organization wallet.

**Does the refusal come back as an HTTP error?** On `/mcp` it comes back as HTTP `200` with `isError: true` and a code. Branch on the code.

**What happens to a call that is already running when the cap is reached?** It finishes and is charged at its full cost. Caps are checked when a call arrives.

**Do top-ups expire?** No. Top-ups on CoreSpeed Pro never expire. Signup credits expire 90 days after the grant, and monthly plan credits reset monthly.