Browse documentation
LlamaIndex: unlob as a retriever or MCP tool
A retriever with no index to build, or the MCP tool spec.
As a retriever
The natural fit. Hits are already passages, so there is nothing to chunk, embed or store — the retriever is a function call.
import os, httpx
from llama_index.core.retrievers import BaseRetriever
from llama_index.core.schema import NodeWithScore, TextNode
class UnlobRetriever(BaseRetriever):
def __init__(self, top_k: int = 8, **kwargs):
self.top_k = top_k
self._client = httpx.Client(
base_url="https://api.unlob.com",
headers={"x-api-key": os.environ["UNLOB_API_KEY"]},
timeout=20,
)
super().__init__(**kwargs)
def _retrieve(self, query_bundle) -> list[NodeWithScore]:
r = self._client.get("/search", params={
"q": query_bundle.query_str,
"limit": self.top_k,
# One row per near-duplicate cluster, so the context window does not
# fill with the same syndicated article.
"collapse": "story",
})
r.raise_for_status()
body = r.json()
if body.get("partial"):
# Incomplete, not empty. Worth surfacing rather than silently
# returning three nodes as though that were the whole corpus.
print("unlob: partial results — part of the corpus was unreachable")
return [
NodeWithScore(
node=TextNode(
text=h["snippet"],
id_=h["id"],
metadata={k: h.get(k) for k in
("url", "host", "title", "independent_sources")},
),
score=h["score"],
)
for h in body["results"]
]
from llama_index.core.query_engine import RetrieverQueryEngine
engine = RetrieverQueryEngine.from_args(UnlobRetriever(top_k=8))
print(engine.query("How does tokio schedule tasks?"))
score is comparable within a response, not between responses — do not threshold on an
absolute value. Use min_quality or min_host_rank if you want a floor that means the
same thing every time.
As MCP tools
llama-index-tools-mcp converts the server into a tool spec, which gives an agent all
every tool rather than search alone:
from llama_index.tools.mcp import BasicMCPClient, McpToolSpec
mcp_client = BasicMCPClient(
"https://api.unlob.com/mcp",
headers={"x-api-key": os.environ["UNLOB_API_KEY"]},
)
tools = McpToolSpec(client=mcp_client).to_tool_list()
That is worth it when the agent should be able to corroborate a claim or brief itself on an entity, not only search.
Skipping the pipeline
If the question is “what should I know about X”, assemble_context does the whole
retrieval loop server-side and returns a packed, deduplicated, trust-ranked reading set —
one call instead of retrieve-rerank-synthesise.
pack = client.get("/assemble_context",
params={"q": question, "budget": 4000}).json()
What _retrieve does not yet handle
UnlobRetriever._retrieve above prints a warning on partial and calls
r.raise_for_status() on anything else, which turns a 429 into a generic
HTTPStatusError a query engine has no way to act on. The two causes of a 429 want
different handling upstream — one is worth a retry, one is not:
import time # alongside the httpx import above
if r.status_code == 429:
retry_after = r.headers.get("retry-after")
if retry_after:
time.sleep(int(retry_after))
r = self._client.get(r.request.url, headers=r.request.headers)
else:
raise RuntimeError("unlob: out of monthly credits, not retrying")
r.raise_for_status()
A synchronous retriever inside RetrieverQueryEngine.from_args is a reasonable place to
retry a genuine rate limit once, since the caller is blocked on the query anyway; it is the
wrong place to retry a hard-capped key out of credits, which will not clear until the billing period
rolls over regardless of how many times query() is called. See
Rate limits and credits for 402 — account suspension — as the third
case neither branch above covers.
Why this fits LlamaIndex better than most retrieval APIs
A BaseRetriever subclass is meant to return nodes ready to synthesise from, and most
retrieval APIs make that awkward: a raw vector store returns chunks that still need
metadata joined on, or a generic search API returns full pages that then need the actual
relevant span pulled out before synthesis is worth running at all. UnlobRetriever above
skips both steps because snippet already is the span, not a page to re-chunk, and the
metadata LlamaIndex wants attached to each TextNode — url, host, independent_sources
— arrives on the hit directly rather than needing a second lookup. That is also why there is
no embedding model configured anywhere in this retriever: LlamaIndex’s usual job of turning
documents into vectors happens once, upstream, inside unlob’s own index, not per query in
your process.
The one place this differs from a VectorStoreIndex-backed retriever worth calling out
explicitly: top_k here does not trade off against index freshness the way it does with a
local index you built yourself. A larger top_k costs nothing except more tokens in the
synthesis step — the search itself is one credit either way — so the number worth tuning
is how much your synthesiser can usefully read, not how expensive retrieval is to run
against a bigger collection. collapse=story in the params above is what keeps a larger
top_k from spending that budget on five copies of the same wire report instead of five
distinct passages.