# Build an AI Lead Generation Agent With OpenAI Agents SDK and Zenrows

> Most lead agents only rank what a database already holds. This one reads a bot-protected directory live, pulls 78 companies off one page, enriches each from its own site, and scores it against your ICP.

Source: https://www.zenrows.com/blog/ai-lead-generation-agent-openai-agents-sdk-zenrows

This agent reads a bot-protected business directory, pulls 78 companies off a single page, and scores them against your ICP at under 20 seconds each. It uses the OpenAI Agents SDK for the agent loop and [Zenrows Fetch](/products/fetch) for the part that usually breaks: getting the page at all.

The OpenAI SDK gives you an agent loop and a `@function_tool` decorator that turns any Python function into a tool the agent can call. That makes it possible to work from live sources rather than a commercial data provider, whose main drawback is that your agent returns the list every other agent is already querying.

Web scraping for lead generation gets around that, because niche directories, industry-specific business directories, and regional business registries surface leads commercial databases miss. The catch is that lead scraping from those sources is only as good as your access to them. Accessing them programmatically is the hard part, because many of them sit behind bot protection such as [Akamai](/blog/bypass-akamai), [Cloudflare](/blog/bypass-cloudflare), or [DataDome](/blog/selenium-datadome).

Without reliable fetching infrastructure, your agent returns zero leads the moment it meets a protected page. Zenrows Fetch with `mode=auto` handles that in a single call, covering protected access, network selection, and dynamic website support, and returns the page as Markdown your model can read.

This article shows you how to build an AI lead generation agent around Zenrows Fetch, using four tools to fetch a directory, extract prospects, enrich each prospect, and score them against your ICP (ideal customer profile).

## Why Database-Only Lead Agents Hit a Ceiling

A database-only agent can only rank what the provider has already indexed. That creates three constraints:

- **Recall:** The dataset is a snapshot of the provider's corpus, so two agents running against it with similar ICPs produce the same candidate set. You can refine the scoring logic and add qualification criteria, but you cannot out-prompt a source constraint when the provider decides recall.
- **Fixed filters:** Indexed datasets overrepresent companies with structured digital data, leaving niche businesses and operators underrepresented. For those that are represented, you can filter only on defined fields like headcount and revenue. Anything not structurally captured never enters your pipeline.
- **Refresh latency:** A company's headcount stays stable for months, but a hiring surge or funding event opens a sales window that lasts weeks. By the time an indexed dataset catches the change, the signal is stale and the prospect has been contacted.

![Database-only lead agents narrow every company in your market through recall, fixed filters and refresh latency, while live sources return every company the directory lists, read at run time](/blog/_img/ai-lead-generation-agent-database-vs-live-sources.png)

Live sources remove those constraints and introduce a different one. Protected sites block standard HTTP clients and headless browsers, and they rarely do it with an error. A [Cloudflare-protected page can return a 200 status code](/blog/bypass-cloudflare) with a challenge page in the body, so your agent treats the challenge and the empty lead list as valid content.

![A protected directory can answer HTTP 200 OK with a Just a moment challenge page in the body, so the agent extracts no leads; Zenrows Fetch with mode=auto returns the real page as Markdown so extract_leads gets the listings](/blog/_img/ai-lead-generation-agent-silent-block.png)

Zenrows Fetch with `mode=auto` gives you access to protected targets in a single call, at a 99.9 percent average workflow success rate. Wrap it with the SDK's function-tool decorator and your agent gets a `fetch_page` tool that reliably returns web data. If you would rather not write the wrapper yourself, the [Zenrows MCP server](https://docs.zenrows.com/mcp/overview) exposes the same capability to any MCP-compatible agent, and our [walkthrough of the MCP server](/blog/mcp-server) covers the setup.

## What You Will Build

An agent that takes a plain-language ICP description, fetches leads from a protected directory with Zenrows Fetch, enriches each lead from its own website, scores the results against the ICP, and returns ranked JSON.

### Prerequisites

- Python 3.10 or above
- An OpenAI API key, used by the agent and by the extraction and scoring tools. Get one from your [OpenAI developer dashboard](https://platform.openai.com/api-keys)
- A Zenrows API key to retrieve the full contents of web pages. Create one on the [signup page](https://app.zenrows.com/register)

Install the OpenAI Agents SDK, `requests` for the Zenrows calls, `pydantic` for the tool schemas, and `python-dotenv` to load your keys:

```bash
python3 -m pip install openai-agents openai requests pydantic python-dotenv
```

Create a `.env` file in your project root with both keys, and add `.env` to `.gitignore` so the credentials never reach version control.

All code in this tutorial is on [GitHub](https://github.com/ZenRows/ai-lead-generation-agent-openai-agents-sdk-zenrows).

## The Four-Tool Architecture

The work splits into four jobs. Separating them keeps retries local: if enrichment fails on one lead, the agent retries that tool instead of restarting the run. It also keeps each job's output small enough to hand to the next.

| Tool | Input | Output | Why it is separate |
| --- | --- | --- | --- |
| `fetch_page` | a URL | the page as Markdown | the only job that touches the network, so retries and errors are handled in one place |
| `extract_leads` | Markdown plus the source URL | `Lead` records | chunked, because a full directory page exceeds one reliable extraction call |
| `enrich_lead` | a company and its domain | lead enrichment signals from its own site | one fetch per nav link found, so it is the cost driver and worth isolating |
| `score_lead` | enriched lead plus the ICP | a score from 0 to 100 and reasoning | pure reasoning, no network, so it can be re-run without re-fetching |

![The four-tool architecture of the lead generation agent: an ICP prompt goes to the agent, discover_leads runs fetch_page to get the directory as Markdown and extract_leads to turn it into Lead records, then qualify_lead runs enrich_lead and score_lead for each lead, and the agent returns a ranked JSON lead list](/blog/_img/ai-lead-generation-agent-four-tool-architecture.png)

The agent sees only two of them. `discover_leads` runs `fetch_page` and `extract_leads`; `qualify_lead` runs `enrich_lead` and `score_lead`. Passing an enriched lead between two separate tools would mean the agent reproduces every page excerpt as a tool argument, which overflows the context window on a large directory.

## Building the Zenrows Fetch Tool

The fetch tool calls Zenrows Fetch and returns Markdown, with failures returned as text rather than raised, so one bad URL does not end the run and the model can read whether a retry is worth it.

`mode=auto` means Zenrows selects the access configuration each target needs. Simple pages resolve on the lightest path and only protected targets escalate.

Every code block from here lives in one file, `tools.py`. Append each one to the same file.

```python

import os
import re
import requests
from urllib.parse import urlparse, parse_qs, unquote

from agents import function_tool
from dotenv import load_dotenv
from openai import OpenAI
from pydantic import BaseModel

load_dotenv()

ZENROWS_API_KEY = os.getenv("ZENROWS_API_KEY")
ZENROWS_ENDPOINT = "https://api.zenrows.com/v1/"

if not ZENROWS_API_KEY:
    raise RuntimeError("set ZENROWS_API_KEY in your .env file")

client = OpenAI()


# ===========================================================================
# job 1: fetch a page, protected or not
# ===========================================================================

def _fetch_page(url: str) -> str:
    """fetch logic, importable and testable without the agent."""
    params = {
        "url": url,
        "apikey": ZENROWS_API_KEY,
        "mode": "auto",              # start cheap, escalate only when the target needs it
        # markdown strips nav and boilerplate before the model reads it
        "response_type": "markdown"
    }

    try:
        response = requests.get(ZENROWS_ENDPOINT, params=params, timeout=90)
        response.raise_for_status()
    # return failures as text instead of raising, so one bad url does not end the run.
    # each message tells the model whether the failure is worth retrying
    except requests.exceptions.Timeout:
        return "FETCH_ERROR: timed out after 90s. retry once, then move on."
    except requests.exceptions.HTTPError as exc:
        status = exc.response.status_code
        if status in (401, 403):
            return f"FETCH_ERROR: {status}. api key rejected, do not retry."
        if status == 429:
            return "FETCH_ERROR: 429 concurrency limit. wait, then retry."
        if status in (400, 404):
            return f"FETCH_ERROR: {status}. this url is not retrievable, do not retry."
        return f"FETCH_ERROR: {status}. retry once, then report the failure."
    except requests.exceptions.RequestException as exc:
        return f"FETCH_ERROR: {exc}. do not retry."

    content = response.text.strip()
    if not content:
        return "FETCH_ERROR: empty response body. do not retry."

    return content


# the decorator replaces the function with a FunctionTool object, which is not
# callable. keeping the logic in _fetch_page above leaves it testable
@function_tool
def fetch_page(url: str) -> str:
    """Fetch any web page as markdown, including pages behind anti-bot protection.

    Args:
        url: full url of the page to fetch
    """
    return _fetch_page(url)


# ===========================================================================
# job 2: turn a directory page into leads
# ===========================================================================

# matches the tracking links directories wrap around outbound urls
REDIRECT_LINK = re.compile(r"https?://[a-z0-9.-]*/redirect\?[^\s\)\"]+")

# a full directory page is too long for one reliable extraction call
CHUNK_SIZE = 40000

# slices cut mid-listing, and the extraction prompt is told to skip partial
# entries, so a listing straddling a boundary is dropped by both slices.
# overlapping the window carries each boundary listing whole into one of them;
# _key dedupes the entries the overlap sees twice
CHUNK_OVERLAP = 2000
```

`response_type=markdown` gives the model cleaner input for extraction, especially on pages heavy with navigation and boilerplate. If your directory is an API-like endpoint, use `response_type=json` and skip the conversion.

## Extracting the Leads

The returned Markdown goes to `extract_leads`, which parses it into structured records. `discover_leads` at the bottom combines fetching and extraction, because the output of `fetch_page` can exceed the context length on its own.

```python
# strict tool schemas reject bare dicts, so every tool input and output is a model
class Lead(BaseModel):
    company: str
    name: str
    website: str
    source_url: str


class LeadList(BaseModel):
    leads: list[Lead]


def _unwrap_redirects(markdown: str) -> str:
    """rewrite directory tracking links to the destination domain."""
    def replace(match):
        url = match.group(0)
        # the real destination sits url-encoded in the u parameter
        target = parse_qs(urlparse(url).query).get("u", [""])[0]
        if not target:
            return url
        parsed = urlparse(unquote(target))
        return f"{parsed.scheme}://{parsed.netloc}"

    return REDIRECT_LINK.sub(replace, markdown)


def _reject_source_domain(leads: list[dict], source_url: str) -> list[dict]:
    """blank any website that points back at the directory itself."""
    directory = urlparse(source_url).netloc.replace("www.", "")
    for lead in leads:
        host = urlparse(lead["website"]).netloc.replace("www.", "")
        if host == directory:
            lead["website"] = ""
    return leads


def _key(lead: dict) -> str:
    """stable identity for dedupe across chunks."""
    # the model returns the same domain with and without www, so normalise before comparing
    site = lead["website"].lower()
    for prefix in ("https://", "http://", "www."):
        site = site.replace(prefix, "")
    site = site.rstrip("/")
    # fall back to company name so leads without a domain are not collapsed into one
    return site or lead["company"].strip().lower()


def _extract_chunk(chunk: str, source_url: str) -> list[dict]:
    """run one extraction call over a slice of the page."""
    response = client.responses.parse(
        model="gpt-4o-mini",
        max_output_tokens=8000,
        input=[
            {
                "role": "system",
                "content": (
                    "extract every company listed in this page fragment. "
                    "website is the company's own domain, taken from that listing's "
                    "visit website link. "
                    "never use a link from the directory's own domain as the website. "
                    "name is a person's name and is usually absent from a directory "
                    "listing, so leave it empty unless a person is actually named. "
                    "use an empty string for any field the fragment does not state, "
                    "and never invent a value. "
                    "the fragment may start or end mid listing, so skip any partial entry."
                )
            },
            {
                "role": "user",
                "content": f"source_url: {source_url}\n\n{chunk}"
            }
        ],
        text_format=LeadList
    )
    return [lead.model_dump() for lead in response.output_parsed.leads]


def _extract_leads(content: str, source_url: str) -> list[dict]:
    """parse markdown into lead dicts, chunking long pages."""
    # a fetch failure is a string, not a page. parsing it would return an empty
    # list that looks identical to a genuinely empty directory
    if content.startswith("FETCH_ERROR"):
        return []

    # resolve tracking links before the model reads them
    content = _unwrap_redirects(content)

    leads, seen = [], set()
    for start in range(0, len(content), CHUNK_SIZE - CHUNK_OVERLAP):
        chunk = content[start:start + CHUNK_SIZE]
        for lead in _extract_chunk(chunk, source_url):
            key = _key(lead)
            if key in seen:
                continue
            seen.add(key)
            leads.append(lead)

    return _reject_source_domain(leads, source_url)


@function_tool
def extract_leads(content: str, source_url: str) -> list[Lead]:
    """Parse a fetched page into a list of leads.

    Args:
        content: markdown returned by fetch_page
        source_url: url the content came from, recorded on every lead
    """
    return [Lead(**lead) for lead in _extract_leads(content, source_url)]

# fetch and extract are one job from the agent's point of view. exposing them
# separately means the whole page passes through the model's context to get
# from one tool to the next, which overflows the window on a large directory
@function_tool
def discover_leads(source_url: str) -> list[Lead]:
    """Fetch a directory page and return the companies listed on it.

    Args:
        source_url: url of the directory or listing page
    """
    markdown = _fetch_page(source_url)
    if markdown.startswith("FETCH_ERROR"):
        return []
    return [Lead(**lead) for lead in _extract_leads(markdown, source_url)]

# ===========================================================================
# job 3: enrich each lead from its own website
# ===========================================================================
```

A separate test script calls these directly, without the agent:

```python
# test_fetch_extract.py
import json
import os
from tools import _fetch_page, _unwrap_redirects, _extract_leads

SOURCE_URL = "https://clutch.co/it-services"
FIXTURE = "fixtures/directory_page.md"
LEADS_CACHE = "fixtures/leads.json"

# ---- job 1: fetch ---------------------------------------------------------
# cache the page so extraction can be tuned without paying for a fetch each time
if os.path.exists(FIXTURE):
    with open(FIXTURE) as f:
        markdown = f.read()
    print(f"using cached page, {len(markdown)} chars")
else:
    markdown = _fetch_page(SOURCE_URL)
    if markdown.startswith("FETCH_ERROR"):
        raise SystemExit(markdown)
    os.makedirs("fixtures", exist_ok=True)
    with open(FIXTURE, "w") as f:
        f.write(markdown)
    print(f"fetched and cached, {len(markdown)} chars")

unwrapped = _unwrap_redirects(markdown)
print(f"{len(unwrapped)} chars after unwrapping redirects, "
      f"{len(markdown) - len(unwrapped)} saved\n")

# ---- job 2: extract -------------------------------------------------------
leads = _extract_leads(markdown, SOURCE_URL)

with open(LEADS_CACHE, "w") as f:
    json.dump(leads, f, indent=2)

print(f"{len(leads)} leads extracted\n")
for lead in leads[:5]:
    print(f"{lead['company']:<32} {lead['website'] or '(none)'}")

missing_site = sum(1 for lead in leads if not lead["website"])
print(f"\n{missing_site} of {len(leads)} leads have no website")
```

Run it from the project root:

```bash
python3 test_fetch_extract.py
```

Output from a local run against an IT services directory:

```text
using cached page, 543637 chars
457647 chars after unwrapping redirects, 85990 saved

78 leads extracted

Infracore                        https://infracore.net
Miles IT                         https://www.milesit.com
Geniusee                         https://geniusee.com
Peeklogic, LLC                   https://www.peeklogic.com
TechQuarter LLC                  https://www.techquarter.io

0 of 78 leads have no website
```

Counts vary between runs, because directory pages change and the model can read the same page differently.

## Building the Lead Enrichment and Scoring Tools

An extracted lead carries only a company name and a domain, which is not enough to judge fit. Lead enrichment is the step that fixes that, and here it runs against the company's own website rather than a third-party database. The signals that matter live on the company's own site: hiring pages, services pages, case studies, contact pages.

`enrich_lead` reads the homepage, finds the navigation links the site actually exposes, and fetches only those. Guessing paths like `/careers` costs a fetch per miss. Zenrows Fetch is doing the work of a lead enrichment API here, with the difference that the signals come from the live site rather than a vendor's snapshot of it.

![enrich_lead reads Infracore's homepage navigation and follows only the links it has, contact, services, portfolio, about and hiring, never a guessed /careers path, then score_lead scores the lead 80 after 12.8 seconds](/blog/_img/ai-lead-generation-agent-enrich-from-own-site.png)

```python
# matches any markdown link, used to read a company's own navigation
MARKDOWN_LINK = re.compile(r"\[([^\]]+)\]\((https?://[^\s\)]+)\)")

# guessing paths costs a fetch per miss, so read the site's nav instead and
# follow only the links it actually has
SIGNAL_KEYWORDS = {
    "hiring": ["career", "job", "join", "hiring", "work with us", "we are hiring"],
    "services": ["service", "what we do", "solutions", "expertise", "capabilities"],
    "portfolio": ["portfolio", "case stud", "our work", "projects", "clients"],
    "about": ["about", "who we are", "our story", "team"],
    "contact": ["contact", "get in touch", "book a call", "let's talk", "talk to us"],
}


# dict[str, X] is rejected by strict schemas too, which is why signals is a list
class Signal(BaseModel):
    name: str
    found: bool
    url: str
    excerpt: str


class EnrichedLead(BaseModel):
    company: str
    website: str
    signals: list[Signal]


def _discover_links(markdown: str, website: str) -> dict:
    """find real urls for each signal by reading the homepage nav."""
    host = urlparse(website).netloc.replace("www.", "")
    homepage_path = urlparse(website).path.rstrip("/") or "/"
    found = {}

    for text, url in MARKDOWN_LINK.findall(markdown):
        parsed = urlparse(url)

        # only follow links on the company's own domain
        if parsed.netloc.replace("www.", "") != host:
            continue

        # an anchor on the homepage is content already in the homepage excerpt
        if parsed.fragment and (parsed.path.rstrip("/") or "/") == homepage_path:
            continue

        # drop the fragment, it never changes what the server returns
        clean_url = f"{parsed.scheme}://{parsed.netloc}{parsed.path}"

        haystack = f"{text} {url}".lower()
        for signal, keywords in SIGNAL_KEYWORDS.items():
            if signal not in found and any(k in haystack for k in keywords):
                found[signal] = clean_url

    return found


def _enrich_lead(company: str, website: str) -> EnrichedLead:
    """collect signals from the company's own site."""
    # a lead with no domain still flows through to scoring, judged on what little is known
    if not website:
        return EnrichedLead(company=company, website="", signals=[])

    homepage = _fetch_page(website)
    if homepage.startswith("FETCH_ERROR"):
        return EnrichedLead(company=company, website=website, signals=[])

    # the homepage is the most informative single page, and on a one-page site
    # it holds everything the nav links point at
    signals = [Signal(name="homepage", found=True, url=website, excerpt=homepage[:6000])]

    for name, url in _discover_links(homepage, website).items():
        content = _fetch_page(url)
        ok = not content.startswith("FETCH_ERROR")
        signals.append(Signal(
            name=name,
            found=ok,
            url=url,
            # keep excerpts small, scoring reads all of them in one call
            excerpt=content[:2000] if ok else ""
        ))

    return EnrichedLead(company=company, website=website, signals=signals)


@function_tool
def enrich_lead(company: str, website: str) -> EnrichedLead:
    """Collect signals from a company's own website.

    Args:
        company: company name from the directory listing
        website: company's own domain
    """
    return _enrich_lead(company, website)


# ===========================================================================
# job 4: score each lead against the icp
# ===========================================================================
```

Markdown works well when the model does the reading. If you want structured fields instead, Zenrows [Extract](/products/extract) returns JSON without a selector map.

With the signals collected, `score_lead` reads them alongside the ICP description and returns a score from 0 to 100. The instruction keeps the score tied to evidence: a lead with only a domain scores lower than one with a careers page and named case studies.

```python
class Score(BaseModel):
    score: int
    reasoning: str


class ScoredLead(BaseModel):
    company: str
    website: str
    score: int
    reasoning: str


def _score_lead(lead: EnrichedLead, icp: str) -> ScoredLead:
    """score an enriched lead against the icp description."""
    # no fetching happens here. every page was already retrieved in job 3
    summary = "\n\n".join(
        f"{s.name} page: {'found at ' + s.url if s.found else 'not found'}\n{s.excerpt}"
        for s in lead.signals
    )

    response = client.responses.parse(
        model="gpt-4o-mini",
        max_output_tokens=500,
        input=[
            {
                "role": "system",
                "content": (
                    "score how well this company matches the ideal customer profile. "
                    "return an integer from 0 to 100 and one sentence of reasoning. "
                    "base the score only on evidence in the signals provided. "
                    "a missing page is weak evidence, not disqualifying. "
                    "if the signals are too thin to judge, score below 30 and say so."
                )
            },
            {
                "role": "user",
                "content": (
                    f"ideal customer profile:\n{icp}\n\n"
                    f"company: {lead.company}\n"
                    f"website: {lead.website}\n\n"
                    f"signals:\n{summary}"
                )
            }
        ],
        text_format=Score
    )

    return ScoredLead(
        company=lead.company,
        website=lead.website,
        score=response.output_parsed.score,
        reasoning=response.output_parsed.reasoning
    )


@function_tool
def score_lead(lead: EnrichedLead, icp: str) -> ScoredLead:
    """Score an enriched lead against an ICP description.

    Args:
        lead: enriched lead returned by enrich_lead
        icp: plain-language description of the ideal customer
    """
    return _score_lead(lead, icp)

# an EnrichedLead carries page excerpts the model has no reason to read. passing
# it between two tools means the agent has to reproduce all of it as an argument,
# and it summarises instead, so scoring reads a summary rather than the pages
@function_tool
def qualify_lead(company: str, website: str, icp: str) -> ScoredLead:
    """Enrich a lead from its own website and score it against an ICP.

    Args:
        company: company name from the directory listing
        website: company's own domain
        icp: plain-language description of the ideal customer
    """
    enriched = _enrich_lead(company, website)
    return _score_lead(enriched, icp)
```

`qualify_lead` pairs enrichment and scoring for the same reason `discover_leads` pairs fetching and extraction: the bulky intermediate stays inside one Python call.

Run both halves before wiring the agent, so any failure you see later belongs to the agent loop rather than the tools:

```python
# test_enrich_score.py

import json
import os
import time
from tools import _fetch_page, _unwrap_redirects, _extract_leads, _enrich_lead, _score_lead

SOURCE_URL = "https://clutch.co/it-services"
FIXTURE = "fixtures/directory_page.md"
LEADS_CACHE = "fixtures/leads.json"

ICP = (
    "IT services agencies that build custom software for B2B clients, "
    "publish detailed case studies with named clients, and offer ai and cloud services"
)

# ---- job 1: fetch ---------------------------------------------------------
if os.path.exists(FIXTURE):
    with open(FIXTURE) as f:
        markdown = f.read()
    print(f"using cached page, {len(markdown)} chars")
else:
    markdown = _fetch_page(SOURCE_URL)
    if markdown.startswith("FETCH_ERROR"):
        raise SystemExit(markdown)
    os.makedirs("fixtures", exist_ok=True)
    with open(FIXTURE, "w") as f:
        f.write(markdown)
    print(f"fetched and cached, {len(markdown)} chars")

unwrapped = _unwrap_redirects(markdown)
print(f"{len(unwrapped)} chars after unwrapping redirects, "
      f"{len(markdown) - len(unwrapped)} saved\n")

# ---- job 2: extract -------------------------------------------------------
if os.path.exists(LEADS_CACHE):
    with open(LEADS_CACHE) as f:
        leads = json.load(f)
    print(f"using {len(leads)} cached leads\n")
else:
    leads = _extract_leads(markdown, SOURCE_URL)
    with open(LEADS_CACHE, "w") as f:
        json.dump(leads, f, indent=2)
    print(f"{len(leads)} leads extracted\n")

for lead in leads[:5]:
    print(f"{lead['company']:<32} {lead['website'] or '(none)'}")

missing_site = sum(1 for lead in leads if not lead["website"])
print(f"\n{missing_site} of {len(leads)} leads have no website\n")
print("-" * 60 + "\n")

# ---- jobs 3 and 4: enrich and score ---------------------------------------
# enrichment is one fetch per nav link found, so start with a few leads
sample = [lead for lead in leads if lead["website"]][:3]

for lead in sample:
    start = time.time()
    enriched = _enrich_lead(lead["company"], lead["website"])
    elapsed = time.time() - start

    found = [s.name for s in enriched.signals if s.found]
    print(f"{enriched.company}  ({elapsed:.1f}s)")
    print(f"  found: {', '.join(found) if found else 'nothing'}")

    scored = _score_lead(enriched, ICP)
    print(f"  score: {scored.score}")
    print(f"  {scored.reasoning}\n")
```

Run it:

```bash
python3 test_enrich_score.py
```

Output, showing the first two leads:

```text
Infracore  (12.8s)
  found: homepage, contact, services, portfolio, about, hiring
  score: 80
  Infracore fits well within the ideal customer profile as they offer IT services, including cloud solutions and have case studies showcasing customized work for various industries, but they may lack detailed named client presentations.

Miles IT  (12.1s)
  found: homepage, contact, about, services, portfolio, hiring
  score: 75
  Miles IT provides software development and AI services, aligning well with the ideal customer profile, but lacks detailed case studies with named clients which slightly diminishes the score.
```

The page list is what each site exposes in its navigation, so a lead with no careers page has no hiring signal.

## Wiring the Agent and Running It

The agent gets the two composed tools and a fixed sequence in the prompt. A small model is enough here, because the planning is already decided and the real reasoning happens inside `score_lead`.

```python
import asyncio

from agents import Agent, Runner, set_tracing_disabled

from tools import discover_leads, qualify_lead

set_tracing_disabled(True)

SOURCE_URL = "https://clutch.co/it-services"

ICP = (
    "IT services agencies that build custom software for B2B clients, "
    "publish detailed case studies with named clients, and offer ai and cloud services"
)

INSTRUCTIONS = """You find and qualify sales leads from web directories.

Follow this sequence exactly:

1. call discover_leads on the source url the user gives you
2. for each lead that has a website, call qualify_lead with its company, website,
   and the icp description from the user message
3. return the scored leads as a json array sorted by score descending

rules:
- process only the first 10 leads that have a website, then stop and return them
- skip leads with no website, they cannot be scored fairly
- if a tool returns a string starting with FETCH_ERROR, follow the instruction in
  that message. do not retry a call the message tells you not to retry
- never invent a lead, a website, or a score. every value comes from a tool result
"""

agent = Agent(
    name="lead generation agent",
    instructions=INSTRUCTIONS,
    model="gpt-4o-mini",  # the reasoning here is scoring, not planning
    tools=[discover_leads, qualify_lead],
)


async def main():
    result = await Runner.run(
        agent,
        input=f"source url: {SOURCE_URL}\n\nideal customer profile:\n{ICP}",
        max_turns=50,  # one qualify call per lead, plus discovery and the reply
    )

    # the trace shows which tools ran and in what order
    for item in result.new_items:
        if item.type == "tool_call_item":
            print(item.raw_item.name)

    print()
    print(result.final_output)


if __name__ == "__main__":
    asyncio.run(main())
```

Run the agent:

```bash
python3 agent.py
```

The agent prints the ranked leads as JSON. The first two:

```json
[
    {
        "company": "Geniusee",
        "website": "https://geniusee.com",
        "score": 85,
        "reasoning": "Geniusee offers a range of IT services including custom software development, AI solutions, and cloud services, which align closely with the ideal customer profile; however, the absence of detailed case studies with named clients slightly limits the fit."
    },
    {
        "company": "Sigli",
        "website": "https://www.sigli.com",
        "score": 85,
        "reasoning": "Sigli aligns closely with the ideal customer profile by offering custom software development for B2B clients, specializing in AI and cloud services, and publishing case studies, although the specificity of named clients could not be verified."
    }
]
```

Every value comes from a tool result rather than the model's own knowledge, so the tool-call trace is the first thing to check when a run misbehaves.

## Scaling With Batch for Large Source Lists

One directory page rarely produces enough leads to matter, and a directory usually has many. Running each page as its own Fetch call means managing the concurrency and the retries yourself.

![Directory pages one to N go into one Batch job through POST /jobs, which handles the queue, concurrency and retries; each page comes back as Markdown through a result_url and into extract_leads](/blog/_img/ai-lead-generation-agent-scale-with-batch.png)

[Zenrows Batch](/products/batch) takes the whole list as a single managed job and handles the queue, concurrency, retries, and result delivery. The script below lives in `batch.py`:

```python
# batch.py
# Scale discovery across several directory pages with Zenrows Batch.
#
# One Fetch call retrieves one page. Batch takes the whole list as a single
# managed job and handles the queue, concurrency and retries, so a directory
# with many pages costs one submission instead of one call per page.

import time

import requests
from agents import function_tool

from tools import ZENROWS_API_KEY, Lead, _extract_leads

BATCH_ENDPOINT = "https://async.api.zenrows.com/v1/jobs"

# batch authenticates by header, unlike fetch which takes apikey as a query param
BATCH_HEADERS = {"X-API-Key": ZENROWS_API_KEY, "Content-Type": "application/json"}

# a job accepts up to 100,000 urls
MAX_URLS_PER_JOB = 100_000


def _submit_batch(urls: list[str]) -> str:
    """submit a url list as one job, return the job id."""
    if len(urls) > MAX_URLS_PER_JOB:
        raise ValueError(f"a job accepts at most {MAX_URLS_PER_JOB} urls")

    payload = {
        "type": "regular",
        "status": "closed",  # run once, accept no further tasks
        "zenrows_params": {"mode": "auto", "response_type": "markdown"},
        "tasks": [{"url": url} for url in urls],
    }
    response = requests.post(
        BATCH_ENDPOINT, headers=BATCH_HEADERS, json=payload, timeout=30
    )
    response.raise_for_status()
    return response.json()["job_id"]


def _collect_batch(job_id: str, max_attempts: int = 60) -> list[tuple[str, str]]:
    """wait for the job to reach a terminal state, then pull each task's markdown."""
    # terminal run states are completed, stopped and deleted. "failed" is a task
    # status, not a run status, so polling for it never returns
    for _ in range(max_attempts):
        time.sleep(5)
        status = requests.get(
            f"{BATCH_ENDPOINT}/{job_id}", headers=BATCH_HEADERS, timeout=30
        )
        status.raise_for_status()
        run = status.json()["latest_run"]
        if run["status"] in ("completed", "stopped", "deleted"):
            break
    else:
        raise TimeoutError(f"job {job_id} did not reach a terminal state in time")

    response = requests.get(
        f"{BATCH_ENDPOINT}/{job_id}/results", headers=BATCH_HEADERS, timeout=30
    )
    response.raise_for_status()

    pages = []
    for task in response.json()["results"]:
        if task["status"] != "successful":
            continue
        # result_url is presigned and valid for 2 hours; re-list the results for a
        # fresh link rather than storing this one
        markdown = requests.get(task["result_url"], timeout=60).text
        pages.append((task["url"], markdown))

    return pages


@function_tool
def discover_leads_batch(source_urls: list[str]) -> list[Lead]:
    """Fetch several directory pages as one job and return every company listed.

    Args:
        source_urls: directory or listing page urls to fetch
    """
    job_id = _submit_batch(source_urls)

    leads = []
    for url, markdown in _collect_batch(job_id):
        leads.extend(Lead(**lead) for lead in _extract_leads(markdown, url))

    return leads


if __name__ == "__main__":
    # two pages of the same directory, submitted as one job
    urls = [
        "https://clutch.co/it-services",
        "https://clutch.co/it-services?page=2",
    ]
    job_id = _submit_batch(urls)
    print(f"submitted {job_id}")

    for url, markdown in _collect_batch(job_id):
        print(f"{url}: {len(markdown)} chars")
```

Batch is also reachable through the Python SDK, the `zenrows batch` [CLI](https://docs.zenrows.com/cli/batch/introduction), and the [Zenrows MCP server](https://docs.zenrows.com/mcp/overview), none of which need the polling loop. Raw REST keeps the example explicit inside a tutorial about writing tools.

For a full implementation guide, see [running large-scale scraping jobs with Batch](/blog/large-scale-web-scraping-zenrows-batch).

## What to Do With the Leads

This tutorial covers discovery and lead qualification. Outreach is a separate layer and beyond its scope.

In a real workflow, put a human review checkpoint between scoring and action, then push qualifying leads into a CRM such as HubSpot or Salesforce, feed them to an email sequencing tool, or export to CSV and hand them over.

## 78 Companies From One Directory Page

You now have an AI lead generation agent that reads a directory with Zenrows Fetch, pulls out every company listed, visits each company's domain, and scores them against a plain-language ICP. The run in this tutorial pulled 78 companies off a single page. `agent.py` scores the first 10 of them per run, at under 20 seconds each; raise or remove that cap in `INSTRUCTIONS` for a full pass.

That is automated lead generation with no vendor list underneath it: the agent reads the source, and the source is whatever directory your market actually uses.

To productionise it, wrap `agent.py` in an HTTP endpoint and call it from n8n, Zapier, or Power Automate.

The complete code is on [GitHub](https://github.com/ZenRows/ai-lead-generation-agent-openai-agents-sdk-zenrows).

## FAQs

### What Is the Difference Between This Agent and Using Apollo or Clay?

Indexed datasets return the same companies when everyone queries with the same filters, and the filters are limited to what the dataset captures. This agent reads the source directly and scores each company on signals from its own site, so two teams with the same ICP can still surface different leads.

### Does This Agent Work on LinkedIn?

No. LinkedIn's user agreement prohibits automated access, and enforcement can result in account restrictions. Point the agent at sources that permit automated access instead, such as public directories, business registries, and review platforms.

### Can the Agent Replace an SDR?

It replaces the discovery and qualification pass. In the run above, the agent pulled 78 companies off a single page and scored them against a plain-language ICP at under 20 seconds each, 10 per run by default. Outreach stays with a person, and this tutorial stops at the handoff.

### What Is Lead Enrichment?

Lead enrichment is the step that turns a bare company record into something you can judge. A directory gives you a name and a domain; enrichment adds the signals that decide fit, such as whether the company is hiring, what it says it does, and whether it publishes client work.

Most tools do this by matching your record against a vendor database. This agent does it by reading the company's own website, which means the signals are current and are not limited to the fields a provider chose to capture.

### Can AI Agents Do Lead Generation End to End?

Not yet, and the honest boundary is worth drawing. AI agents for lead generation are reliable at the mechanical parts: finding sources, pulling records, enriching them, and scoring them against a profile. This agent does all four. The judgment calls, deciding whether a 70 is worth an email and what that email says, stay with a person.

### How Much Does It Cost to Run This Agent on Zenrows?

A standard request costs one credit, and `mode=auto` selects the configuration each target needs, up to 25 credits where a page requires the full configuration. These weights are fixed across all plans.

In this tutorial the directory page resolved at the standard rate. Enrichment is the cost driver, because it fetches the homepage plus each nav link it finds, which ran to five or six fetches per lead. A 10-lead run therefore costs roughly 60 credits. You only pay for successful requests, so a blocked page costs nothing. See the [Zenrows pricing documentation](https://docs.zenrows.com/first-steps/pricing) for details.

### Can I Use This With a Different LLM?

Partly. The Zenrows tools are plain Python functions, so the agent wrapper can use any model. `extract_leads` and `score_lead` depend on OpenAI's structured output API, so those two calls need replacing if you switch providers.

### How Do I Add This Agent to a Workflow Automation Tool Like n8n, Zapier, or Power Automate?

Expose `agent.py` through an HTTP endpoint and trigger it as a webhook step. Pass the source URL and the ICP in the request body.
