HTTP 200 with isError true: the most common MCP integration bug
A CoreSpeed tool call can fail inside an HTTP 200. How the two error layers differ, which codes to retry, and which ones need a person instead.

An MCP tool call can fail in two places. Transport and identity failures arrive as HTTP 4xx or 5xx, and the tool never ran. Tool failures arrive as HTTP 200 with result.isError: true, because the request was authorized and then the tool refused or failed. Code that treats every 200 as success misses the second layer entirely.
| Aspect | HTTP 4xx or 5xx | HTTP 200 with isError: true |
|---|---|---|
| What failed | transport or identity | the tool refused or failed |
| Did the tool run | no | it was authorized and dispatched |
| Where the code is | error.code, with type and message | structuredContent.error.code |
| Examples | 401 invalid credential, 403 agent suspended, 404 unknown path, 503 key authority unreachable | payment_required, key_spend_limit_exceeded, needs_reauth, jwt_session_required, an approval hold |
| Retry | 429 after Retry-After; 500 and 503 with backoff; the rest never | rate_limited, upstream_unreachable, internal once; holds and caps never |
Why does a failed tool call come back as HTTP 200?
MCP is JSON-RPC carried over HTTP. The HTTP status describes the carrier: did the request reach the server, and was the credential accepted. Once that is settled, the server answers the JSON-RPC request. The tool's own outcome lives inside that answer, and the protocol marks a refusal or failure with isError: true on the result.
CoreSpeed follows the protocol. A billing hold, an API key past its monthly cap, a suspended organization, an account that needs reauthorization, and a session-only tool called with an API key all arrive the same way: HTTP 200, isError: true, and a code to branch on.
{
"result": {
"isError": true,
"content": [
{ "type": "text", "text": "Payment required. Metered calls are paused until the organization adds credit." }
],
"structuredContent": { "error": { "code": "payment_required" } }
}
}
The text is written for the agent to read. structuredContent.error.code is the value your code should branch on.
What are the two layers?
The first layer is transport and identity. Every route answers a 4xx or 5xx with a nested error object carrying type, code and message, and the tool never ran. A 401 means the credential itself is wrong: a missing header, a revoked key, an expired token. A 403 means the key is valid but its agent is suspended. A 404 endpoint_not_found means an unknown path, or a connector this environment does not offer. A 503 means the key authority was unreachable, and CoreSpeed failed closed rather than trust an unverified key.
The second layer is tool execution. The request was authorized. Then the tool refused, the upstream call failed, or a policy held it. Everything in this layer is HTTP 200.
So the first branch in your client is result.isError, before you parse anything else. The second branch is the code. The full list of codes is in the error reference, and the operational decisions are on the error handling page.
Which codes should you retry?
| Code | Layer | Retry? |
|---|---|---|
401, 402, 403, 404 | HTTP | Never. Nothing changes without an action by the client, an admin, or support. |
429 | HTTP | After Retry-After. |
503, 500 | HTTP | With backoff. Keep error.request_id from a 500 for support. |
rate_limited | Tool | After retry_after_seconds. |
upstream_unreachable, token_refresh_failed, platform_db_unavailable | Tool | With backoff. On a write, check whether the change landed before repeating it. |
internal | Tool | Once. |
payment_required, key_spend_limit_exceeded, org_suspended | Tool | Never. A person adds credit, raises the cap, or contacts support. |
needs_reauth, jwt_session_required, an approval hold | Tool | Never. The state does not clear on retry. |
The write rule deserves emphasis. After upstream_unreachable, upstream_response_invalid or cancelled on a call that creates or changes something, the change may have landed. Read it back before you send it again.
Which errors need a person rather than a retry?
payment_required means the organization is past its billing threshold. Later metered calls are refused before the tool runs. Discovery with tools/list, account reads and key management keep working through the hold. An org admin adds credit and the hold clears. key_spend_limit_exceeded means one API key reached its monthly cap; other keys are unaffected, and the cap is raised or the month rolls over. org_suspended is administrative, and adding credit does not clear it.
Caps are checked at the request boundary. An action already in flight can finish above the remaining amount. So your code should expect a refusal on the next request, never an abort in the middle of one. The details are on the billing page.
needs_reauth means an upstream grant can no longer be used. The connector's tools stay in tools/list, and calls fail until the member reauthorizes in Dashboard → Connectors. Do not rotate the CoreSpeed key. Do not rewrite client configuration. Neither one is broken.
jwt_session_required means the tool runs only under a member session. The manage__keys_*, manage__agents_*, manage__accounts_*, manage__whoami and manage__switch_org tools are listed for an API key but refuse it. Switch credential; a retry with the same key answers the same way.
ambiguous_account means several accounts of one connector are in scope. Pass account with one of the aliases in structuredContent.error.aliases. not_configured means this environment lacks an app registration the connector needs. Reconnecting does not help; that one is for support.
An approval hold is the one error the agent should wait on. The held call returns at once with an error-shaped receipt that names an approval id. Call manage__approval_wait with that id. It blocks until a person decides and returns the original tool's result, or a denial or expiry error. Do not re-send the call: a fresh tools/call is a new request and can raise a second card.
How does the cs CLI model these decisions?
The cs CLI collapses the same codes into exit codes, and the grouping is a good model for any client:
0: ok1: usage error or refused3: not signed in4: the account is unable to act right now (no organization, payment required, suspended, admin needed)5: unreachable or rate-limited6: waiting for a human's approval
Each group is one move. Fix the credential. Fix the account. Wait and retry. Wait for a person. A client that maps every code to one of those four moves handles the whole surface without special cases.
FAQ
Is isError: true a JSON-RPC error?
No. A JSON-RPC error means the request itself was malformed or could not be dispatched. isError is a successful dispatch whose tool refused or failed.
Should my client retry any 200 at all?
Only the tool-level codes in the table: rate_limited, upstream_unreachable, token_refresh_failed, platform_db_unavailable, and internal once. Everything else needs a change by a person first.
Where can I see the failure afterward? Dashboard → Activity records the outcome, the code and the text the agent saw, bounded to 512 characters. Prompts and request content are never copied into the trail.
What if a write failed with upstream_unreachable?
Treat it as unknown. Read the resource back through the same connector before repeating the write.