unlob Docs
Browse documentation

LlamaIndex: unlob as a retriever or MCP tool

A retriever with no index to build, or the MCP tool spec.

As a retriever

The natural fit. Hits are already passages, so there is nothing to chunk, embed or store — the retriever is a function call.

import os, httpx
from llama_index.core.retrievers import BaseRetriever
from llama_index.core.schema import NodeWithScore, TextNode

class UnlobRetriever(BaseRetriever):
    def __init__(self, top_k: int = 8, **kwargs):
        self.top_k = top_k
        self._client = httpx.Client(
            base_url="https://api.unlob.com",
            headers={"x-api-key": os.environ["UNLOB_API_KEY"]},
            timeout=20,
        )
        super().__init__(**kwargs)

    def _retrieve(self, query_bundle) -> list[NodeWithScore]:
        r = self._client.get("/search", params={
            "q": query_bundle.query_str,
            "limit": self.top_k,
            # One row per near-duplicate cluster, so the context window does not
            # fill with the same syndicated article.
            "collapse": "story",
        })
        r.raise_for_status()
        body = r.json()
        if body.get("partial"):
            # Incomplete, not empty. Worth surfacing rather than silently
            # returning three nodes as though that were the whole corpus.
            print("unlob: partial results — part of the corpus was unreachable")
        return [
            NodeWithScore(
                node=TextNode(
                    text=h["snippet"],
                    id_=h["id"],
                    metadata={k: h.get(k) for k in
                              ("url", "host", "title", "independent_sources")},
                ),
                score=h["score"],
            )
            for h in body["results"]
        ]
from llama_index.core.query_engine import RetrieverQueryEngine
engine = RetrieverQueryEngine.from_args(UnlobRetriever(top_k=8))
print(engine.query("How does tokio schedule tasks?"))

score is comparable within a response, not between responses — do not threshold on an absolute value. Use min_quality or min_host_rank if you want a floor that means the same thing every time.

As MCP tools

llama-index-tools-mcp converts the server into a tool spec, which gives an agent all every tool rather than search alone:

from llama_index.tools.mcp import BasicMCPClient, McpToolSpec

mcp_client = BasicMCPClient(
    "https://api.unlob.com/mcp",
    headers={"x-api-key": os.environ["UNLOB_API_KEY"]},
)
tools = McpToolSpec(client=mcp_client).to_tool_list()

That is worth it when the agent should be able to corroborate a claim or brief itself on an entity, not only search.

Skipping the pipeline

If the question is “what should I know about X”, assemble_context does the whole retrieval loop server-side and returns a packed, deduplicated, trust-ranked reading set — one call instead of retrieve-rerank-synthesise.

pack = client.get("/assemble_context",
                  params={"q": question, "budget": 4000}).json()

What _retrieve does not yet handle

UnlobRetriever._retrieve above prints a warning on partial and calls r.raise_for_status() on anything else, which turns a 429 into a generic HTTPStatusError a query engine has no way to act on. The two causes of a 429 want different handling upstream — one is worth a retry, one is not:

import time  # alongside the httpx import above

if r.status_code == 429:
    retry_after = r.headers.get("retry-after")
    if retry_after:
        time.sleep(int(retry_after))
        r = self._client.get(r.request.url, headers=r.request.headers)
    else:
        raise RuntimeError("unlob: out of monthly credits, not retrying")
r.raise_for_status()

A synchronous retriever inside RetrieverQueryEngine.from_args is a reasonable place to retry a genuine rate limit once, since the caller is blocked on the query anyway; it is the wrong place to retry a hard-capped key out of credits, which will not clear until the billing period rolls over regardless of how many times query() is called. See Rate limits and credits for 402 — account suspension — as the third case neither branch above covers.

Why this fits LlamaIndex better than most retrieval APIs

A BaseRetriever subclass is meant to return nodes ready to synthesise from, and most retrieval APIs make that awkward: a raw vector store returns chunks that still need metadata joined on, or a generic search API returns full pages that then need the actual relevant span pulled out before synthesis is worth running at all. UnlobRetriever above skips both steps because snippet already is the span, not a page to re-chunk, and the metadata LlamaIndex wants attached to each TextNode — url, host, independent_sources — arrives on the hit directly rather than needing a second lookup. That is also why there is no embedding model configured anywhere in this retriever: LlamaIndex’s usual job of turning documents into vectors happens once, upstream, inside unlob’s own index, not per query in your process.

The one place this differs from a VectorStoreIndex-backed retriever worth calling out explicitly: top_k here does not trade off against index freshness the way it does with a local index you built yourself. A larger top_k costs nothing except more tokens in the synthesis step — the search itself is one credit either way — so the number worth tuning is how much your synthesiser can usefully read, not how expensive retrieval is to run against a bigger collection. collapse=story in the params above is what keeps a larger top_k from spending that budget on five copies of the same wire report instead of five distinct passages.

Next