Browse documentation
Rate limits and credits
429 and 402 mean opposite things, and one 429 means two different things.
Two independent limits
A per-minute rate limit, set by your plan, applied per key and counted in credits. A call
takes its cost from the minute, so a ground call uses five times what a search does. A call
that costs no credits, such as /account, still takes one. Exceed the limit and you get
429 with a retry-after saying how long the wait actually is. The window is one minute
long and is shared by the whole fleet, so the figure on your plan is the figure you get.
A monthly allowance of credits, also set by your plan. Every call costs a published number of credits, listed on Pricing. What happens when you run out depends on the plan:
- On Free, the key hard-caps: a call that costs more than the credits left gets
429until the period rolls over, and the refusal is not billed. A cheaper call that still fits keeps working, so a key with 3 credits left refusesgroundbut answers a search. There is no card on file, so free means free. - On paid plans, calls continue past the allowance. The pricing page promises that rather than failing closed — an agent in production should not stop working because a counter rolled over.
What a call is billed
- A call that succeeds costs its published credits.
- A call that runs and fails costs 1 credit: a
400, a404, a timeout, a503. - A refused call costs nothing:
401,402, and both kinds of429.
The headers
Every authenticated response carries the rate headers, on success and on 429 alike. Both
count credits:
x-ratelimit-limit: 600
x-ratelimit-remaining: 587
A call that ran also says what it was billed:
x-credits-charged: 5
On a rate-limit 429, one more:
retry-after: 14
retry-after is a real number of seconds, not a constant. It is the time until your
current minute turns over, so it is somewhere between 1 and 60 depending on where in the
window you were refused. Sleep the value you were given — a client that sleeps one second
and retries into the same window is simply refused again.
There is no x-ratelimit-reset. retry-after already carries the wait, so there is no
reset timestamp to look for.
Telling the two 429s apart
This is the part worth reading carefully. Both a rate-limit refusal and a hard-capped key
without the credits for a call return 429, and they call for opposite responses:
| Rate limit | Out of credits, hard-capped | |
|---|---|---|
| Body | rate limit exceeded | not enough credits left this month: this call costs 5 and 3 remain, … or monthly credits used up, … |
retry-after | seconds until the window turns over | absent |
| Retry? | Yes, after retry-after seconds | No. The same call will not fit until the period rolls over. |
So: retry a 429 only when retry-after is present. A retry loop that ignores the
body and hammers a hard-capped key produces thousands of identical failures and never
succeeds.
And 402 is different again
402 means the account is suspended — a subscription was cancelled, or a payment
failed. Retrying changes nothing; somebody has to settle billing in the
console.
Summary:
| Status | Meaning | Retry? |
|---|---|---|
429 with retry-after | Too fast | Yes, after retry-after seconds |
429 without retry-after | Hard-capped, and out of credits for this call | No |
402 | Account suspended | No |
401 | Bad key — or a key too new to have propagated | Once, after a minute |
The published rate is the whole account’s rate
Your per-minute rate is counted once for the key, not once per server answering it. Every
request you make spends from the same minute, whichever machine handles it, so the number on
your plan is the number you can actually rely on — a floor and a ceiling at once. Size your
concurrency against the published figure and pace off x-ratelimit-remaining.
503 at capacity
Separately from your own limits, the service can be too busy to take a request at all. It
says so immediately — 503 with a body of server is at capacity; retry shortly and a
retry-after — rather than holding the request in a queue until it times out. It is a
transient condition and it is worth retrying; see Errors.
Over MCP
Identical. The same key, the same limits, the same headers on the HTTP response, and a tool call costs what its REST operation costs.
One difference in how a refusal reaches you: over MCP, a suspended account or a key out of
credits arrives as a tool error inside the JSON-RPC frame, not as a 402 or 429 status.
That is deliberate — a client reads a transport error on an established session as the server
having gone away and tears the session down, which turns a billing problem into an apparent
outage.
Rate limiting still answers 429 at the transport, because that is what a client’s own
backoff logic understands.
Connecting is free of credits: initialize, notifications/initialized, tools/list and
ping are never billed, and they never touch your credit allowance or your credits-a-minute
window. Free is not unlimited, though — session frames have their own generous per-minute
allowance, so a client that reconnects or re-lists tools in a tight loop can still be told
429. Open a session once and keep using it, and you will never see it.
Unauthenticated calls
The description routes — /openapi.json, /mcp/tools.json, /llms.txt, /skill.md,
/guides/{slug}, /describe — need no key and cost no credits. They are not unlimited: a
flood of unauthenticated requests from one address is refused with 429 and a retry-after,
and so is a flood of calls presenting keys that do not authenticate. Normal reading of the
description routes is nowhere near it.
Staying inside them
- Read
x-ratelimit-remainingand pace yourself, rather than discovering the limit. - Read
x-credits-charged, or pollGET /account, to know where the month stands. - Use
limit=0withfacets=trueto size a topic in one search instead of five exploratory ones. - Use
collapse=storyso you are not spending credits — and context — on the same article repeatedly. assemble_contextis one call in place of a retrieval loop that would be several.
Next
- Errors
- Retries and backoff
- Pricing and plans
- Logging and monitoring calls — alerting on these headers before you hit them