unlob Docs
Browse documentation

Rate limits and credits

429 and 402 mean opposite things, and one 429 means two different things.

Two independent limits

A per-minute rate limit, set by your plan, applied per key and counted in credits. A call takes its cost from the minute, so a ground call uses five times what a search does. A call that costs no credits, such as /account, still takes one. Exceed the limit and you get 429 with a retry-after saying how long the wait actually is. The window is one minute long and is shared by the whole fleet, so the figure on your plan is the figure you get.

A monthly allowance of credits, also set by your plan. Every call costs a published number of credits, listed on Pricing. What happens when you run out depends on the plan:

  • On Free, the key hard-caps: a call that costs more than the credits left gets 429 until the period rolls over, and the refusal is not billed. A cheaper call that still fits keeps working, so a key with 3 credits left refuses ground but answers a search. There is no card on file, so free means free.
  • On paid plans, calls continue past the allowance. The pricing page promises that rather than failing closed — an agent in production should not stop working because a counter rolled over.

What a call is billed

  • A call that succeeds costs its published credits.
  • A call that runs and fails costs 1 credit: a 400, a 404, a timeout, a 503.
  • A refused call costs nothing: 401, 402, and both kinds of 429.

The headers

Every authenticated response carries the rate headers, on success and on 429 alike. Both count credits:

x-ratelimit-limit: 600
x-ratelimit-remaining: 587

A call that ran also says what it was billed:

x-credits-charged: 5

On a rate-limit 429, one more:

retry-after: 14

retry-after is a real number of seconds, not a constant. It is the time until your current minute turns over, so it is somewhere between 1 and 60 depending on where in the window you were refused. Sleep the value you were given — a client that sleeps one second and retries into the same window is simply refused again.

There is no x-ratelimit-reset. retry-after already carries the wait, so there is no reset timestamp to look for.

Telling the two 429s apart

This is the part worth reading carefully. Both a rate-limit refusal and a hard-capped key without the credits for a call return 429, and they call for opposite responses:

Rate limitOut of credits, hard-capped
Bodyrate limit exceedednot enough credits left this month: this call costs 5 and 3 remain, … or monthly credits used up, …
retry-afterseconds until the window turns overabsent
Retry?Yes, after retry-after secondsNo. The same call will not fit until the period rolls over.

So: retry a 429 only when retry-after is present. A retry loop that ignores the body and hammers a hard-capped key produces thousands of identical failures and never succeeds.

And 402 is different again

402 means the account is suspended — a subscription was cancelled, or a payment failed. Retrying changes nothing; somebody has to settle billing in the console.

Summary:

StatusMeaningRetry?
429 with retry-afterToo fastYes, after retry-after seconds
429 without retry-afterHard-capped, and out of credits for this callNo
402Account suspendedNo
401Bad key — or a key too new to have propagatedOnce, after a minute

The published rate is the whole account’s rate

Your per-minute rate is counted once for the key, not once per server answering it. Every request you make spends from the same minute, whichever machine handles it, so the number on your plan is the number you can actually rely on — a floor and a ceiling at once. Size your concurrency against the published figure and pace off x-ratelimit-remaining.

503 at capacity

Separately from your own limits, the service can be too busy to take a request at all. It says so immediately — 503 with a body of server is at capacity; retry shortly and a retry-after — rather than holding the request in a queue until it times out. It is a transient condition and it is worth retrying; see Errors.

Over MCP

Identical. The same key, the same limits, the same headers on the HTTP response, and a tool call costs what its REST operation costs.

One difference in how a refusal reaches you: over MCP, a suspended account or a key out of credits arrives as a tool error inside the JSON-RPC frame, not as a 402 or 429 status. That is deliberate — a client reads a transport error on an established session as the server having gone away and tears the session down, which turns a billing problem into an apparent outage.

Rate limiting still answers 429 at the transport, because that is what a client’s own backoff logic understands.

Connecting is free of credits: initialize, notifications/initialized, tools/list and ping are never billed, and they never touch your credit allowance or your credits-a-minute window. Free is not unlimited, though — session frames have their own generous per-minute allowance, so a client that reconnects or re-lists tools in a tight loop can still be told 429. Open a session once and keep using it, and you will never see it.

Unauthenticated calls

The description routes — /openapi.json, /mcp/tools.json, /llms.txt, /skill.md, /guides/{slug}, /describe — need no key and cost no credits. They are not unlimited: a flood of unauthenticated requests from one address is refused with 429 and a retry-after, and so is a flood of calls presenting keys that do not authenticate. Normal reading of the description routes is nowhere near it.

Staying inside them

  • Read x-ratelimit-remaining and pace yourself, rather than discovering the limit.
  • Read x-credits-charged, or poll GET /account, to know where the month stands.
  • Use limit=0 with facets=true to size a topic in one search instead of five exploratory ones.
  • Use collapse=story so you are not spending credits — and context — on the same article repeatedly.
  • assemble_context is one call in place of a retrieval loop that would be several.

Next