unlob Docs
Browse documentation

Pydantic AI: unlob as a typed toolset

A typed toolset from the MCP server, or a typed tool of your own.

As an MCP server

import os
from pydantic_ai import Agent
from pydantic_ai.mcp import MCPServerStreamableHTTP

unlob = MCPServerStreamableHTTP(
    "https://api.unlob.com/mcp",
    headers={"x-api-key": os.environ["UNLOB_API_KEY"]},
)

agent = Agent(
    "anthropic:claude-opus-5",
    toolsets=[unlob],
    system_prompt=(
        "Search with unlob. Use assemble_context for open questions and web_search "
        "for specific ones. Call corroborate before stating anything as established."
    ),
)

async def main():
    async with agent:
        result = await agent.run("What is corroborated about the Acme acquisition?")
        print(result.output)

async with agent opens the MCP connection for the run and closes it after. Without it, each call reconnects.

A typed tool

Pydantic AI’s argument model becomes the tool schema, which makes constraints — enums, ranges — part of what the model is told rather than something you validate afterwards.

import os, httpx
from typing import Literal
from pydantic import BaseModel, Field
from pydantic_ai import Agent, RunContext

class Hit(BaseModel):
    title: str
    url: str
    text: str
    independent_sources: int = Field(
        0, description="Distinct sites asserting this. 1 means a single source."
    )

class SearchResult(BaseModel):
    incomplete: bool = Field(
        False,
        description="True when part of the corpus was unreachable. The answer is "
                    "short because of a failure, not because there is nothing more.",
    )
    hits: list[Hit]

_client = httpx.AsyncClient(
    base_url="https://api.unlob.com",
    headers={"x-api-key": os.environ["UNLOB_API_KEY"]},
    timeout=20,
)

agent = Agent("anthropic:claude-opus-5")

@agent.tool
async def web_search(
    ctx: RunContext[None],
    query: str,
    mode: Literal["keyword", "semantic", "hybrid"] = "hybrid",
    limit: int = 5,
) -> SearchResult:
    """Search the web for passages. Returns passage text, not links."""
    r = await _client.get("/search", params={
        "q": query,
        "mode": mode,
        "limit": min(limit, 20),
        "min_independent_sources": 2,
        "collapse": "story",
    })
    r.raise_for_status()
    body = r.json()
    return SearchResult(
        incomplete=bool(body.get("partial")),
        hits=[
            Hit(
                title=h["title"],
                url=h["url"],
                text=h["snippet"],
                independent_sources=h.get("independent_sources", 0),
            )
            for h in body["results"]
        ],
    )

Putting incomplete in the return model rather than raising is deliberate: partial is a 200, and an exception would hide a result that is real but short.

Structured research output

Where Pydantic AI earns its keep is the output type. Ask for citations as a schema and you get them as a schema:

class Finding(BaseModel):
    claim: str
    sources: list[str] = Field(description="URLs supporting this claim")
    independently_corroborated: bool

class Research(BaseModel):
    summary: str
    findings: list[Finding]

agent = Agent("anthropic:claude-opus-5", output_type=Research, toolsets=[unlob])

independently_corroborated maps directly onto what corroborate returns, so the model has something real to fill it from rather than a guess.

Modelling the out-of-credits case, not just the incomplete one

SearchResult.incomplete above handles partial: true — a real answer that came back short. A 429 is a different shape of problem: the call did not complete at all, and web_search currently lets r.raise_for_status() turn it into a generic HTTPStatusError. Distinguishing the two 429 causes before that happens gives the agent something it can act on instead of a raw exception:

if r.status_code == 429:
    if "retry-after" in r.headers:
        raise ModelRetry(f"Rate limited, retry after {r.headers['retry-after']}s")
    raise ModelRetry("Monthly credits exhausted for this key — do not retry")
r.raise_for_status()

Pydantic AI’s ModelRetry feeds the message back to the model as the reason the call failed, which is exactly the case where wording matters: “credits exhausted, do not retry” stops the agent from calling the tool again in a loop, where a bare HTTPStatusError gives it no signal either way. 402 — account suspended — is worth its own branch for the same reason; see Rate limits and credits for the full set.

Dependency injection instead of a module-level client

The _client above is defined once at module scope, which is fine for a script but awkward to test: swapping the key or pointing at a different base URL means monkeypatching a global. Pydantic AI’s RunContext and deps_type exist for exactly this case. Passing an httpx.AsyncClient through dependencies rather than importing it directly means a test can construct the agent with a mocked client and never touch a real key or a real network call:

from dataclasses import dataclass

@dataclass
class Deps:
    client: httpx.AsyncClient

agent = Agent("anthropic:claude-opus-5", deps_type=Deps)

@agent.tool
async def web_search(ctx: RunContext[Deps], query: str) -> SearchResult:
    r = await ctx.deps.client.get("/search", params={"q": query})
    ...

ctx.deps.client is then whatever the caller constructed the agent with — the real authenticated client in production, a respx-mocked one in a test suite that asserts on min_independent_sources being set correctly without spending a real request. For anything beyond a one-off script, this is worth the small extra indirection over the module-level client shown further up the page.

The same Deps dataclass is also where a per-tenant key belongs in a multi-tenant service: constructing one Agent per request with that request’s own key in deps, rather than a single shared client, keeps one tenant’s credits and one tenant’s key from ever crossing into another tenant’s run. TestModel from pydantic_ai.models.test pairs naturally with the mocked client above, letting a test assert on which tool the agent called and with which arguments without spending a real model call either, which is usually the faster way to catch a wrong parameter name before it ships than waiting on a live run against the API.

Next