Browse documentation
CrewAI: give a crew unlob search tools
A research tool for a research crew.
A tool
CrewAI tools are classes with a Pydantic argument schema.
import os, httpx
from typing import Type
from crewai.tools import BaseTool
from pydantic import BaseModel, Field
class SearchInput(BaseModel):
query: str = Field(description='Supports AND, OR, -exclude and "exact phrase".')
vertical: str | None = Field(None, description="e.g. code, science. Omit to auto-route.")
limit: int = Field(5, description="Maximum results, 1-20.")
class UnlobSearch(BaseTool):
name: str = "unlob_search"
description: str = (
"Search the web for passages. Returns passage text, not links, so the result "
"is usually the answer rather than something to fetch. Each hit carries "
"independent_sources: how many distinct sites assert it."
)
args_schema: Type[BaseModel] = SearchInput
def _run(self, query: str, vertical: str | None = None, limit: int = 5) -> str:
params = {
"q": query,
"limit": min(limit, 20),
# Applied here rather than left to the agent: drop single-source claims,
# and fold near-duplicates so the crew does not read one article five times.
"min_independent_sources": 2,
"collapse": "story",
}
if vertical:
params["vertical"] = vertical
r = httpx.get(
"https://api.unlob.com/search",
params=params,
headers={"x-api-key": os.environ["UNLOB_API_KEY"]},
timeout=20,
)
if r.status_code == 429 and "retry-after" not in r.headers:
return "Monthly credits exhausted. Do not retry; report this to the user."
r.raise_for_status()
body = r.json()
header = ("WARNING: incomplete results — part of the corpus was unreachable.\n\n"
if body.get("partial") else "")
hits = "\n\n".join(
f"{h['title']} — {h['url']} ({h.get('independent_sources', 0)} independent sources)\n{h['snippet']}"
for h in body["results"]
)
return header + (hits or "No results found.")
A verification tool worth pairing with it
The reason to use unlob in a crew rather than a generic search tool:
class CorroborateInput(BaseModel):
id: str = Field(description="A passage id from a search result.")
class UnlobCorroborate(BaseTool):
name: str = "unlob_corroborate"
description: str = (
"Check whether a claim is independently reported. Given a passage id, returns "
"the distinct hosts carrying that story, grouped and ranked by authority. Use "
"this before stating anything as established fact — forty results echoing one "
"source and four independent reports look identical in a result list."
)
args_schema: Type[BaseModel] = CorroborateInput
def _run(self, id: str) -> str:
r = httpx.get(
"https://api.unlob.com/corroborate",
params={"id": id},
headers={"x-api-key": os.environ["UNLOB_API_KEY"]},
timeout=20,
)
r.raise_for_status()
b = r.json()
hosts = ", ".join(s["host"] for s in b["sources"])
return (f"{b['independent_sources']} independent sources "
f"({b['merged_duplicates']} duplicates folded in): {hosts}")
Using them
from crewai import Agent, Task, Crew
researcher = Agent(
role="Research analyst",
goal="Find what is actually established about {topic}, and what is only asserted",
backstory=(
"You distinguish corroborated fact from repetition. You never state something "
"as established without checking how many independent sources carry it."
),
tools=[UnlobSearch(), UnlobCorroborate()],
)
task = Task(
description="Research {topic}. Corroborate every claim you intend to report.",
expected_output="A brief listing each finding with its independent source count.",
agent=researcher,
)
Crew(agents=[researcher], tasks=[task]).kickoff(inputs={"topic": "…"})
The backstory is doing real work there. A crew given a search tool will search; a crew told what distinguishes a fact from a repetition will use the second tool.
The other failure _run should not swallow
UnlobSearch._run above already branches on a 429 with no retry-after header — the
hard-capped out-of-credits case, where retrying inside the crew’s own retry logic (CrewAI
retries a failing tool by default) would just produce the same failure several times for no
benefit. Worth adding the case it does not yet cover:
if r.status_code == 402:
return "Account suspended — billing needs attention. Do not retry this task."
A 402 is not a credits problem and not a rate limit; it means the account itself stopped
working, and no crew-level retry policy fixes it. Surfacing that distinctly rather than
letting raise_for_status() produce a generic CrewAI tool-execution error means the agent’s
final answer says what actually happened instead of “the search tool failed.”
Passing the crew a floor, not a suggestion
min_independent_sources is set in _run itself, in UnlobSearch above, rather than left
as an argument the LLM fills in. That is the point of writing the tool as a class instead of
handing CrewAI a bare function schema: a value baked into the request the agent cannot see
or lower is a floor, and a value the model is merely told to set is a suggestion it can talk
itself out of under a tight deadline. Do the same with collapse=story for anything that
searches news or syndicated content, where the same wire report otherwise shows up under
five different hostnames and reads to the crew like five sources.
Which agent gets which tool
A crew with several agents does not need every agent holding both tools above. A researcher
that only ever gathers should get UnlobSearch alone; put UnlobCorroborate on a separate
verification-focused agent, or a reviewer step later in the process, rather than trusting
every agent in the crew to remember to call it. That split matches how CrewAI’s own process
model works — a sequential crew can pass the researcher’s findings to a dedicated verifier
task whose entire backstory is about checking rather than finding, which tends to produce
more consistent corroboration than one agent juggling both jobs under time pressure from its
own task description. A hierarchical crew gets the same benefit from a manager agent whose
instructions explicitly require a corroboration check before anything is reported as
established, rather than leaving it to whichever worker agent happens to pick up the claim.
The two tools are deliberately separate classes rather than one tool with a mode
parameter, for exactly this reason: separate tools can be handed to separate agents, and a
single combined tool cannot be split later without changing every agent that already holds
it. Keeping them separate from the start costs nothing when there is only one agent in the
crew, and saves a refactor the day a second one joins, since neither tool’s schema changes
when a new agent starts holding it.