Browse documentation
Pydantic AI: unlob as a typed toolset
A typed toolset from the MCP server, or a typed tool of your own.
As an MCP server
import os
from pydantic_ai import Agent
from pydantic_ai.mcp import MCPServerStreamableHTTP
unlob = MCPServerStreamableHTTP(
"https://api.unlob.com/mcp",
headers={"x-api-key": os.environ["UNLOB_API_KEY"]},
)
agent = Agent(
"anthropic:claude-opus-5",
toolsets=[unlob],
system_prompt=(
"Search with unlob. Use assemble_context for open questions and web_search "
"for specific ones. Call corroborate before stating anything as established."
),
)
async def main():
async with agent:
result = await agent.run("What is corroborated about the Acme acquisition?")
print(result.output)
async with agent opens the MCP connection for the run and closes it after. Without it,
each call reconnects.
A typed tool
Pydantic AI’s argument model becomes the tool schema, which makes constraints — enums, ranges — part of what the model is told rather than something you validate afterwards.
import os, httpx
from typing import Literal
from pydantic import BaseModel, Field
from pydantic_ai import Agent, RunContext
class Hit(BaseModel):
title: str
url: str
text: str
independent_sources: int = Field(
0, description="Distinct sites asserting this. 1 means a single source."
)
class SearchResult(BaseModel):
incomplete: bool = Field(
False,
description="True when part of the corpus was unreachable. The answer is "
"short because of a failure, not because there is nothing more.",
)
hits: list[Hit]
_client = httpx.AsyncClient(
base_url="https://api.unlob.com",
headers={"x-api-key": os.environ["UNLOB_API_KEY"]},
timeout=20,
)
agent = Agent("anthropic:claude-opus-5")
@agent.tool
async def web_search(
ctx: RunContext[None],
query: str,
mode: Literal["keyword", "semantic", "hybrid"] = "hybrid",
limit: int = 5,
) -> SearchResult:
"""Search the web for passages. Returns passage text, not links."""
r = await _client.get("/search", params={
"q": query,
"mode": mode,
"limit": min(limit, 20),
"min_independent_sources": 2,
"collapse": "story",
})
r.raise_for_status()
body = r.json()
return SearchResult(
incomplete=bool(body.get("partial")),
hits=[
Hit(
title=h["title"],
url=h["url"],
text=h["snippet"],
independent_sources=h.get("independent_sources", 0),
)
for h in body["results"]
],
)
Putting incomplete in the return model rather than raising is deliberate: partial is a
200, and an exception would hide a result that is real but short.
Structured research output
Where Pydantic AI earns its keep is the output type. Ask for citations as a schema and you get them as a schema:
class Finding(BaseModel):
claim: str
sources: list[str] = Field(description="URLs supporting this claim")
independently_corroborated: bool
class Research(BaseModel):
summary: str
findings: list[Finding]
agent = Agent("anthropic:claude-opus-5", output_type=Research, toolsets=[unlob])
independently_corroborated maps directly onto what corroborate returns, so the model has
something real to fill it from rather than a guess.
Modelling the out-of-credits case, not just the incomplete one
SearchResult.incomplete above handles partial: true — a real answer that came back
short. A 429 is a different shape of problem: the call did not complete at all, and
web_search currently lets r.raise_for_status() turn it into a generic HTTPStatusError.
Distinguishing the two 429 causes before that happens gives the agent something it can act
on instead of a raw exception:
if r.status_code == 429:
if "retry-after" in r.headers:
raise ModelRetry(f"Rate limited, retry after {r.headers['retry-after']}s")
raise ModelRetry("Monthly credits exhausted for this key — do not retry")
r.raise_for_status()
Pydantic AI’s ModelRetry feeds the message back to the model as the reason the call
failed, which is exactly the case where wording matters: “credits exhausted, do not retry”
stops the agent from calling the tool again in a loop, where a bare HTTPStatusError gives
it no signal either way. 402 — account suspended — is worth its own branch for the same
reason; see Rate limits and credits for the full set.
Dependency injection instead of a module-level client
The _client above is defined once at module scope, which is fine for a script but awkward
to test: swapping the key or pointing at a different base URL means monkeypatching a global.
Pydantic AI’s RunContext and deps_type exist for exactly this case. Passing an
httpx.AsyncClient through dependencies rather than importing it directly means a test can
construct the agent with a mocked client and never touch a real key or a real network call:
from dataclasses import dataclass
@dataclass
class Deps:
client: httpx.AsyncClient
agent = Agent("anthropic:claude-opus-5", deps_type=Deps)
@agent.tool
async def web_search(ctx: RunContext[Deps], query: str) -> SearchResult:
r = await ctx.deps.client.get("/search", params={"q": query})
...
ctx.deps.client is then whatever the caller constructed the agent with — the real
authenticated client in production, a respx-mocked one in a test suite that asserts on
min_independent_sources being set correctly without spending a real request. For anything
beyond a one-off script, this is worth the small extra indirection over the module-level
client shown further up the page.
The same Deps dataclass is also where a per-tenant key belongs in a multi-tenant service:
constructing one Agent per request with that request’s own key in deps, rather than a
single shared client, keeps one tenant’s credits and one tenant’s key from ever crossing into
another tenant’s run. TestModel from pydantic_ai.models.test pairs naturally with the
mocked client above, letting a test assert on which tool the agent called and with which
arguments without spending a real model call either, which is usually the faster way to
catch a wrong parameter name before it ships than waiting on a live run against the API.