Browse documentation
Every HTTP status the unlob API returns
Every status worth handling, and what to do about each one.
Shape
Errors are a plain-text reason; the status code carries the meaning. There is no JSON error envelope to parse, and no error code field — the status and the body text are the contract.
HTTP/1.1 429 Too Many Requests
retry-after: 14
x-ratelimit-limit: 600
x-ratelimit-remaining: 0
rate limit exceeded
The table
| Status | Meaning | Do |
|---|---|---|
400 | The request is malformed — usually a required parameter missing or a value the endpoint cannot use | Fix the call. Never retry. |
401 | No key, an unknown key, a revoked key, or a key too new to have propagated | See below |
402 | Account suspended for billing | Settle it in the console. Never retry. |
404 | No such passage id, or no such guide slug | Not an error to retry — the id does not exist |
408 | The request exceeded the server timeout | Retry once, with a narrower query |
429 with retry-after | Over the per-minute rate, or too many unauthenticated calls from your address | Sleep the retry-after value and retry |
429 without retry-after | A hard-capped key without the credits this call costs | Never retry — it will not change this period |
500 | Something failed server-side, and not because the index was unreachable — that is 503 | Retry once with backoff |
503 with retry-after | The server is at capacity and refused the request before running it | Sleep the retry-after value and retry |
503 without retry-after | No shard could serve the request — the node has none configured, or every one it asked failed | Not yours to fix; retry with backoff |
401 on a brand-new key
The most common surprise. Keys propagate to the serving fleet on a refresh cycle, so a key
minted seconds ago can answer 401 for up to about a minute.
The rule: 401 on a key that has never worked is worth one retry after a minute.
401 on a key that worked yesterday is not — it was revoked, or the account lapsed, and
retrying will not change either. Check the console.
401 deliberately does not distinguish “wrong key” from “revoked key”, so probing tells an
attacker nothing.
The two 429s
Both mean stop, and they mean it for different lengths of time:
if r.status_code == 429:
if "retry-after" in r.headers:
time.sleep(int(r.headers["retry-after"])) # the real wait; then retry
else:
raise CreditsExhausted(r.text) # this month; do not retry
Branching on the header rather than the status is what stops a retry loop hammering a hard-capped key thousands of times to no effect. See Rate limits and credits.
The retry-after on a rate-limit 429 is a real number of seconds — the time until your
minute turns over — not a constant 1. Sleep the value; a fixed one-second sleep retries
into the same window and is refused again. An expensive call on a small plan is told a
correspondingly longer wait: a ten-credit call against thirty credits a minute is told
twenty seconds.
A 429 can also come from an unauthenticated call: the description routes are open and
free, but a flood from one address is refused, as are repeated calls whose key does not
authenticate. It carries a retry-after like any other rate-limit refusal.
Two different 503s
Both are 503, and only the body and the presence of retry-after tell them apart.
| Body | retry-after | Meaning | Do |
|---|---|---|---|
server is at capacity; retry shortly | present | The server was too busy to take the request. It was refused before it ran, rather than queued until it timed out. | Retryable. Sleep the header value, then retry with backoff. |
| a shard-unavailability reason | absent | No shard could serve the request | Retry with backoff, but it will not clear on your schedule — the data plane is unavailable |
The distinction matters for a retry policy: the first is load, clears in seconds, and is the
one to retry aggressively. The second is an outage, and hammering it does not help. A
sustained 503 from a search route is the signal that the data plane, rather than your
request, is the problem.
A request refused at capacity never ran, so it costs no credits. A call that ran and then failed costs 1 credit.
408 — narrow the query, do not just retry
A timeout usually means the query was expensive: a very broad term with no filters, a large
limit, or mode=semantic over a wide vertical. Retrying the identical request is likely
to time out identically.
Add a filter, name the vertical, or lower limit, and it will usually go through.
Successful responses that are not what you wanted
Two cases return 200 and still need handling.
partial: true — part of the corpus was unreachable, so the answer is incomplete. Retry,
or report it as incomplete. Never treat it as evidence of absence. This is the most
important non-error in the API; see Reading a result.
An empty results array — usually genuine. The index is admission-controlled, so it
holds what was judged worth keeping. Confirm with why_not before
concluding the web is silent on a topic.
Over MCP
The split matters, and it is deliberate:
| Situation | How it arrives |
|---|---|
| No key, unknown key | HTTP 401 — a caller with no session has nowhere to put a JSON-RPC error |
| Over the per-minute rate, or over the session-frame allowance | HTTP 429 — this is what a client’s own backoff understands |
| Unknown tool | JSON-RPC error -32601 |
| Bad arguments | JSON-RPC error -32602 |
| Suspended account, out of credits | A tool error in the result, not a status |
| The tool ran and failed | A tool error in the result |
| The server is at capacity | HTTP 503 with retry-after |
A tool error looks like this — note it is a successful JSON-RPC response:
{
"jsonrpc": "2.0",
"id": 3,
"result": {
"content": [{ "type": "text", "text": "account suspended — settle billing in the console" }],
"isError": true
}
}
Billing refusals are tool errors rather than statuses because a client reads a transport error on an established session as the server having gone away, and tears the session down. A billing problem should not present as an outage.