unlob Docs
Browse documentation

Retries and backoff

What to retry, what never to retry, and how fast.

The decision, in one table

StatusRetryHow
429 with retry-afterYesSleep the header value — it is the real wait, 1 to 60 seconds — then retry.
429 without retry-afterNoHard-capped and out of credits for this call. It will not change this period.
402NoSuspended account. Somebody has to settle billing.
401, first ever use of a keyOnceWait a minute — it may not have propagated yet.
401, a key that worked beforeNoRevoked or lapsed. Check the console.
400, 404NoThe request is wrong. Retrying will not fix it.
408OnceAnd narrow the query — see below.
503 with retry-afterYesThe server is at capacity. Sleep the header value, then back off.
500, 503 without retry-afterYesExponential backoff with jitter.
Network error, connection resetYesExponential backoff with jitter.
200 with partial: trueYesTreat like 503: part of the corpus was unreachable.

The one that catches people is 429 meaning two unrelated things. Branch on retry-after, not on the status — a loop that retries every 429 will hammer a hard-capped key thousands of times and never succeed.

503 splits the same way, and for the same reason. With a retry-after the server was at capacity and refused the request before running it: transient, and worth retrying on the schedule the header names. Without one, no shard could serve the request — retry with backoff, but a sustained run of them is an outage rather than something your client can wait out. The body text says which; see Errors.

Every read is idempotent

Every endpoint is a GET and nothing has side effects, so a retry cannot duplicate anything. There is no idempotency key to manage, and none is needed.

The only cost of a retry is that the retried call is billed again. That matters for the retries you should not be making: a retry loop against a hard cap spends nothing (a refusal is not billed), but a retry loop against 500 costs a credit a time.

A retry that behaves

import random, time, httpx

RETRYABLE = {408, 500, 502, 503, 504}

def search(client: httpx.Client, params: dict, attempts: int = 4) -> dict:
    for attempt in range(attempts):
        r = client.get("/search", params=params)

        if r.status_code == 429:
            after = r.headers.get("retry-after")
            if after is None:
                # Hard-capped key out of credits. Retrying changes nothing.
                raise RuntimeError(f"out of credits: {r.text}")
            # Sleep what the server said, never a fixed second: this is the time until
            # the window turns over, and a shorter sleep retries into the same window.
            time.sleep(int(after))
            continue

        if r.status_code == 503 and "retry-after" in r.headers:
            # At capacity — the request never ran. The header is the wait.
            time.sleep(int(r.headers["retry-after"]))
            continue

        if r.status_code in RETRYABLE and attempt < attempts - 1:
            # Full jitter: without it, every client that failed together retries
            # together, and the recovering service gets the same spike again.
            time.sleep(random.uniform(0, 2 ** attempt))
            continue

        r.raise_for_status()
        body = r.json()

        if body.get("partial") and attempt < attempts - 1:
            # 200, but part of the corpus was unreachable. A short answer here is a
            # symptom, not a finding — retrying usually gets the whole thing.
            time.sleep(random.uniform(0, 2 ** attempt))
            continue

        return body

    raise RuntimeError("exhausted retries")

Three things that example does deliberately:

  • Jitter, not a fixed doubling. Clients that failed together will otherwise retry together and re-create the spike that broke the thing they are waiting for.
  • Retries partial. It is a 200, so raise_for_status will not catch it, and a caller that ignores it silently reports an incomplete answer as a complete one.
  • Raises on a hard cap rather than looping. Failing loudly at the point of the decision beats a retry budget draining into nothing.
  • Honours retry-after wherever it appears, on a 429 and on an at-capacity 503 alike, instead of substituting its own guess.

Pacing before you get refused

Backoff is what you do after being refused; pacing is how you avoid it.

Every authenticated response carries x-ratelimit-limit and x-ratelimit-remaining. Read them and slow down as remaining falls, rather than sprinting into a 429. There is no x-ratelimit-reset, and none is needed: the window is one minute long, and a refusal tells you exactly how much of it is left in retry-after.

Pace against the published rate as written. It is counted once for your key across the whole service, not once per server answering you, so there is no aggregate headroom above it to discover.

Spending fewer requests helps more than any retry policy:

  • limit=0&facets=true sizes a topic in one call, instead of five exploratory searches.
  • collapse=story stops you paying for the same article five times.
  • assemble_context is one call in place of a whole retrieval loop.

408: retry differently, not just again

A timeout usually means the query was expensive — a very broad term with no filters, a large limit, or semantic mode across a wide vertical. The identical request will probably time out identically.

Name the vertical, add a filter, or lower limit, then retry. That is a retry with new information, which is the only kind worth making twice.

Over MCP

Most clients handle transport-level retries themselves. Two things to know:

  • 429 still arrives as an HTTP status, because that is what client backoff understands. Session frames — initialize, notifications/initialized, tools/list, ping — cost no credits, but they are not unlimited; a client that re-lists tools in a tight loop can be refused one.
  • A suspended account or a key out of credits arrives as a tool error, not a status — your retry logic will not see it, and the agent reading the result has to. Do not let a tool error labelled isError disappear silently into a loop.

The transport is stateless, so a retried request needs no session recovery: any replica can serve it, and the key is re-presented every time.

Next