Build a web research multi-agent system with AG2 and Zenrows
In an AG2 group chat every tool response enters the shared history. Register Zenrows Fetch as a typed tool so clean content, not a failed fetch, is what circulates.
A web research pipeline in AG2 divides the work across a shared context. It only produces useful output when the researcher can actually retrieve the page it was given. In AG2 every tool response enters a shared history that every agent reads on its turn, so a failed fetch does not stay isolated. Each agent that follows reads and forwards that failure, spending tokens on turns that had nothing real to work with.
AG2 is the actively maintained community fork of Microsoft's AutoGen, with native MCP client support added in v0.12. This tutorial uses the OpenAI client.
Registering Zenrows Fetch as a typed AG2 tool, with Adaptive Stealth Mode on, gives the researcher a fetch layer that works, so clean content enters the history from the first turn. Zenrows is web data infrastructure built for the dynamic and protected pages a standard HTTP client cannot reliably reach.
What follows builds a researcher, analyst, and critic pipeline from scratch, and compares token usage across two runs to show what changes when the fetch layer works.

Prerequisites
- Python 3.10 or later.
- AG2 v0.12.2.
- A Zenrows account and API key.
- An OpenAI account and API key.
The full code is on GitHub.
Set up an AG2 group chat
Install AG2 with its OpenAI client:
pip install "ag2[openai]==0.12.2"
Two things to know before you run anything.
First, the package name and the import name differ. You install with pip install "ag2[openai]" and import with import autogen. Using import ag2 raises a ModuleNotFoundError in v0.12.2. AG2 is deprecating the classic framework at v1.0, at which point import ag2 becomes correct.
Second, ag2 does not bundle the OpenAI client, which is why the [openai] extra is required.
Load your API keys from a .env file:
import os
from dotenv import load_dotenv
import autogen
from autogen import AssistantAgent, GroupChat, GroupChatManager, UserProxyAgent
# load api keys from .env
load_dotenv()
llm_config = {
"config_list": [
{
"model": "gpt-4o-mini",
"api_key": os.getenv("OPENAI_API_KEY"),
}
],
# temperature 0 for deterministic output
"temperature": 0,
}
Next, define the four agents:
# user_proxy executes tools and terminates on TERMINATE
user_proxy = UserProxyAgent(
name="user_proxy",
human_input_mode="NEVER",
code_execution_config=False,
is_termination_msg=lambda msg: "TERMINATE" in (msg.get("content") or ""),
)
# researcher fetches and summarizes web content
researcher = AssistantAgent(
name="researcher",
system_message=(
"You are a research agent. When given a URL, fetch its web content and "
"summarize what you find. Be factual and note anything that looks "
"incomplete or missing from the page. Pass your summary to the analyst."
),
llm_config=llm_config,
)
# analyst extracts structured fields from researcher output
analyst = AssistantAgent(
name="analyst",
system_message=(
"You are a data analyst. Extract structured fields, such as plan names "
"and prices, from the content the researcher provides. Present the "
"result as a clean list. If the researcher's content has no usable "
"data, say so explicitly instead of guessing."
),
llm_config=llm_config,
)
# critic validates analyst output and terminates with TERMINATE
critic = AssistantAgent(
name="critic",
system_message=(
"You are a critic. Validate the analyst's structured output against "
"the researcher's original content. Flag any gaps, missing fields, or "
"inconsistencies. When the output checks out or no further progress is "
"possible, say so and end your message with TERMINATE."
),
llm_config=llm_config,
)
user_proxy sits in the group chat as the tool execution agent. It runs any registered function the LLM requests and ends the chat when an agent message contains TERMINATE.
Wire them into a group chat and run it:
# wire agents into a group chat capped at 6 rounds
group_chat = GroupChat(
agents=[user_proxy, researcher, analyst, critic],
messages=[],
max_round=6,
)
# manager routes messages between agents
manager = GroupChatManager(
groupchat=group_chat,
llm_config=llm_config,
)
# kick off the chat with a fetch-and-extract task
chat_result = user_proxy.initiate_chat(
manager,
message=(
"Fetch the product listing from this page and extract the "
"product name and price: https://www.walmart.com/ip/AirPods-Pro-3/17835006350"
),
)
GroupChat caps turns with max_round. If you are coming from two-party chats where initiate_chat accepts max_turns, note that the parameter name is different here.
For reliable token numbers, call autogen.gather_usage_summary() on the agent list. chat_result.cost reports zero for group chats and is not a usable source:
# gather token usage across all agents including the manager
usage_summary = autogen.gather_usage_summary(
[user_proxy, researcher, analyst, critic, manager]
)
print(usage_summary)
Notice that the researcher's system message tells it to fetch web content, but nothing is registered that can do that. The model interprets the instruction as best it can and admits it cannot reach the URL. This is where the pipeline breaks.
Here is what the baseline run produces:
# output truncated
Fetch the product listing from this page and extract the product name and
price: https://www.walmart.com/ip/AirPods-Pro-3/17835006350
--------------------------------------------------------------------------------
Next speaker: researcher
researcher (to chat_manager):
I'm unable to access external URLs directly to fetch content. However, you can
visit the provided link to view the product listing for the AirPods Pro 3 on
Walmart's website.
--------------------------------------------------------------------------------
Next speaker: analyst
analyst (to chat_manager):
I cannot access external URLs or fetch content from them. Please provide the
product name and price directly, and I can help you structure that information.
--------------------------------------------------------------------------------
Next speaker: critic
critic (to chat_manager):
The analyst's output is consistent with the researcher's original content. Both
parties acknowledge the inability to access external URLs. There are no gaps or
missing fields, as the responses align in their limitations. TERMINATE
===== USAGE SUMMARY (ALL AGENTS) =====
{'usage_including_cached_inference': {'total_cost': 0.00025245,
'gpt-4o-mini-2024-07-18': {'cost': 0.00025245, 'prompt_tokens': 759,
'completion_tokens': 231, 'total_tokens': 990}}}
| Run | Prompt tokens | Completion tokens | Total | Cost |
|---|---|---|---|---|
| Baseline (no fetch tool) | 759 | 231 | 990 | $0.000252 |
The researcher had no fetch tool, so it never reached the page. The analyst had nothing to work with, and the critic ended the chat cleanly once both agents confirmed the same limitation. Registering a typed fetch tool at the source is what changes this.
Register Zenrows as an AG2 tool
To give the researcher access to JavaScript-rendered and anti-bot-protected pages, register Zenrows Fetch as a typed tool. AG2 builds the tool schema from your function's type hints, so annotating the parameters is not optional.
Add your Zenrows API key to .env:
ZENROWS_API_KEY=your_zenrows_api_key_here
Then define the fetch function:
import requests
from typing import Annotated
from autogen.tools import Tool
ZENROWS_API_KEY = os.getenv("ZENROWS_API_KEY")
ZENROWS_ENDPOINT = "https://api.zenrows.com/v1/"
def fetch_page_content(
url: Annotated[str, "The target URL to fetch through Zenrows using adaptive stealth mode."],
) -> str:
"""Fetch a URL through Zenrows Fetch and return clean Markdown."""
params = {
"url": url,
"apikey": ZENROWS_API_KEY,
"mode": "auto",
"response_type": "markdown",
}
response = requests.get(ZENROWS_ENDPOINT, params=params, timeout=60)
response.raise_for_status()
return response.text
mode=auto is Adaptive Stealth Mode. Zenrows selects the configuration each page needs, escalating to JavaScript rendering or premium proxies only when the target requires it, so you do not have to guess the combination. response_type=markdown returns clean Markdown, ready to enter the group chat history without inflating the token count with raw HTML.
Wrap the function in an AG2 Tool and register it with both agents:
# wrap the fetch function as an ag2 tool
fetch_tool = Tool(
name="fetch_page_content",
description=(
"Fetch a URL through Zenrows and return clean Markdown content from "
"JavaScript-rendered and anti-bot-protected pages."
),
func_or_tool=fetch_page_content,
)
# expose the tool schema to the researcher
fetch_tool.register_for_llm(researcher)
# user_proxy executes the tool when the researcher calls it
fetch_tool.register_for_execution(user_proxy)
register_for_llm exposes the schema to the researcher so the LLM knows the tool exists and when to call it. register_for_execution tells AG2 which agent runs the function when the LLM requests it. In a group chat that is always the UserProxyAgent.
Update the researcher's system message to name the tool:
# update researcher to explicitly use the fetch tool
researcher = AssistantAgent(
name="researcher",
system_message=(
"You are a research agent. When given a URL, use the fetch_page_content "
"tool to retrieve its web content, then summarize what you find. Be "
"factual and note anything that looks incomplete or missing from the "
"page. Pass your summary to the analyst."
),
llm_config=llm_config,
)
Here is the full pipeline with Zenrows integrated:
import os
from typing import Annotated
import requests
from dotenv import load_dotenv
import autogen
from autogen import AssistantAgent, GroupChat, GroupChatManager, UserProxyAgent
from autogen.tools import Tool
load_dotenv()
ZENROWS_API_KEY = os.getenv("ZENROWS_API_KEY")
ZENROWS_ENDPOINT = "https://api.zenrows.com/v1/"
llm_config = {
"config_list": [
{
"model": "gpt-4o-mini",
"api_key": os.getenv("OPENAI_API_KEY"),
}
],
"temperature": 0,
}
# user_proxy executes tools and terminates on TERMINATE
user_proxy = UserProxyAgent(
name="user_proxy",
human_input_mode="NEVER",
code_execution_config=False,
is_termination_msg=lambda msg: "TERMINATE" in (msg.get("content") or ""),
)
# researcher is told to use the fetch tool by name
researcher = AssistantAgent(
name="researcher",
system_message=(
"You are a research agent. When given a URL, use the fetch_page_content "
"tool to retrieve its web content, then summarize what you find. Be "
"factual and note anything that looks incomplete or missing from the "
"page. Pass your summary to the analyst."
),
llm_config=llm_config,
)
# analyst extracts structured fields from researcher output
analyst = AssistantAgent(
name="analyst",
system_message=(
"You are a data analyst. Extract structured fields, such as plan names "
"and prices, from the content the researcher provides. Present the "
"result as a clean list. If the researcher's content has no usable "
"data, say so explicitly instead of guessing."
),
llm_config=llm_config,
)
# critic validates analyst output and terminates with TERMINATE
critic = AssistantAgent(
name="critic",
system_message=(
"You are a critic. Validate the analyst's structured output against "
"the researcher's original content. Flag any gaps, missing fields, or "
"inconsistencies. When the output checks out or no further progress is "
"possible, say so and end your message with TERMINATE."
),
llm_config=llm_config,
)
def fetch_page_content(
url: Annotated[str, "The target URL to fetch through Zenrows using adaptive stealth mode."],
) -> str:
"""Fetch a URL through Zenrows Fetch and return clean Markdown."""
params = {
"url": url,
"apikey": ZENROWS_API_KEY,
"mode": "auto",
"response_type": "markdown",
}
response = requests.get(ZENROWS_ENDPOINT, params=params, timeout=60)
response.raise_for_status()
return response.text
# wrap the fetch function as an ag2 tool
fetch_tool = Tool(
name="fetch_page_content",
description=(
"Fetch a URL through Zenrows and return clean Markdown content from "
"JavaScript-rendered and anti-bot-protected pages."
),
func_or_tool=fetch_page_content,
)
# expose the tool schema to the researcher
fetch_tool.register_for_llm(researcher)
# user_proxy executes the tool when the researcher calls it
fetch_tool.register_for_execution(user_proxy)
# wire agents into a group chat capped at 6 rounds
group_chat = GroupChat(
agents=[user_proxy, researcher, analyst, critic],
messages=[],
max_round=6,
)
# manager routes messages between agents
manager = GroupChatManager(
groupchat=group_chat,
llm_config=llm_config,
)
# kick off the chat with a fetch-and-extract task
chat_result = user_proxy.initiate_chat(
manager,
message=(
"Fetch the product listing from this page and extract the "
"product name and price: https://www.walmart.com/ip/AirPods-Pro-3/17835006350"
),
)
print("\n===== FULL CHAT HISTORY =====")
for msg in chat_result.chat_history:
speaker = msg.get("name") or msg.get("role")
print(f"[{speaker}]: {msg.get('content')}\n")
print("\n===== USAGE SUMMARY (ALL AGENTS) =====")
# gather token usage across all agents including the manager
usage_summary = autogen.gather_usage_summary(
[user_proxy, researcher, analyst, critic, manager]
)
print(usage_summary)

Run the pipeline and compare token spend
The researcher calls fetch_page_content, Zenrows retrieves the page through Adaptive Stealth Mode, and clean Markdown enters the shared history. The analyst extracts the product name and price, and the critic validates the output and terminates cleanly.
# output truncated
Fetch the product listing from this page and extract the product name and
price: https://www.walmart.com/ip/AirPods-Pro-3/17835006350
--------------------------------------------------------------------------------
Next speaker: researcher
researcher (to chat_manager):
***** Suggested tool call: fetch_page_content *****
Arguments: {"url":"https://www.walmart.com/ip/AirPods-Pro-3/17835006350"}
>>>>>>>> EXECUTING FUNCTION fetch_page_content...
[Markdown content returned from the Walmart page]
--------------------------------------------------------------------------------
Next speaker: analyst
analyst (to chat_manager):
- Product Name: Apple AirPods Pro 3
- Price: $189.99 (was $249.00, you save $59.01)
--------------------------------------------------------------------------------
Next speaker: critic
critic (to chat_manager):
The analyst's structured output is consistent with the researcher's original
content. The product name "Apple AirPods Pro 3" and the price "$189.99 (was
$249.00, you save $59.01)" accurately reflect the information extracted from
the page. No gaps, missing fields, or inconsistencies were found.
TERMINATE
===== USAGE SUMMARY (ALL AGENTS) =====
{'usage_including_cached_inference': {'total_cost': 0.0011466,
'gpt-4o-mini-2024-07-18': {'cost': 0.0011466, 'prompt_tokens': 7100,
'completion_tokens': 136, 'total_tokens': 7236}}}
These are single runs per configuration, on gpt-4o-mini at temperature 0. Absolute numbers vary between runs. What the table shows is which configuration extracted data and which did not.
| Run | Products extracted | Total tokens | Cost |
|---|---|---|---|
| Baseline (no fetch tool) | No | 990 | $0.000252 |
| Fetch with mode=auto | Yes | 7,236 | $0.001147 |
The baseline produced nothing. The working configuration cost about one tenth of a cent.
The Zenrows run spends more tokens because the Markdown from a fully fetched page is far larger than the researcher's brief refusal in the baseline. mode=auto retrieves the real page content on the first call, which gives the analyst accurate product data to structure and the critic real output to validate. Both runs terminate on TERMINATE, so neither wastes tokens on empty turns.
Keep costs predictable in production
Content volume in the shared history is the primary cost driver. Every agent turn after a fetch re-reads and forwards that content, so page size multiplies across the pipeline. max_round sets a ceiling on your worst case, not a cost lever. The baseline terminated in four turns and still cost the least, because it had nothing to circulate. Start with a conservative round limit and raise it only when your pipeline consistently needs more turns to finish naturally.
How many rounds you need depends largely on how reliable the fetch layer is. Every failed or empty fetch that enters the history costs tokens on every subsequent turn, pushing the round count higher than it needs to be.
Fetch reliability is what keeps the round count down. Zenrows runs at a 99.93% success rate across supported targets, so the failed-fetch path stays rare. mode=auto in fetch_page_content handles this at the source, selecting the configuration each page needs and keeping the history clean from the first turn.
Zenrows credit usage with mode=auto scales with what each request needs. A basic request on a static page costs 1 credit. JavaScript rendering costs 5, premium proxies cost 10, and a heavily protected site that needs both costs 25. You are billed once, for the configuration that succeeds, so the internal attempts before it are not charged. The Adaptive Stealth Mode pricing table has the full breakdown.
In a pipeline that fetches the same URL on every run, those credits add up. Caching the Zenrows response after the first call and returning it on later calls within the same session cuts both token spend and credit usage.
For pipelines that research several URLs in one run, Zenrows Batch takes the full list as one managed job. It handles concurrency and retries, so your agents get clean content for every URL without you managing fetch queues.
Wrapping up
AG2's group chat history is what makes multi-agent pipelines powerful, and it is also what makes a bad fetch expensive. Every tool response enters the shared history, and every agent that speaks after reads it and forwards it to the LLM. A fetch that fails or returns nothing useful still costs tokens on every turn that reads it.
Registering Zenrows as a typed fetch tool with mode=auto addresses that at the source. Zenrows decides what the target needs, whether that is plain rendering, JavaScript execution, or premium proxies, without you configuring it by hand. Clean Markdown enters the history on the first call, so the analyst and critic get real content to work with and the pipeline terminates naturally instead of burning its round limit on empty turns.
The full code for both baseline.py and zenrows_pipeline.py is on GitHub. If you are working in a different framework, the same fetch pattern applies to smolagents and the OpenAI Agents SDK.
Frequently asked questions
Do I install autogen or ag2?
Install ag2 from PyPI with pip install "ag2[openai]". It imports as autogen in v0.12.2, so your code uses import autogen. AG2 is deprecating the classic framework at v1.0, at which point import ag2 becomes correct.
AG2 says my tool schema is invalid. Why?
Type hints are required for AG2 to generate the schema. Annotate every input parameter as Annotated[type, description] and give the function a return type. This tutorial uses both: url: Annotated[str, "..."] as the input and -> str as the return. Without them AG2 cannot build the schema and registration fails.
My group chat produces generic output instead of real content. What should I check?
If early turns contain empty or placeholder content, that content flows through every later turn, costing tokens without moving the pipeline forward. Add mode=auto to your Zenrows call so it can escalate to a more capable configuration when the target needs one, rather than defaulting to a basic request that returns an empty or generic response.
Can I connect Zenrows via MCP instead of a registered function?
Yes. AG2 v0.12 added native support for the MCP client. The registered-function approach used here gives more control over the tool schema and error handling, but if you prefer MCP, the Zenrows MCP server covers that setup.
How do I see token usage per turn?
Use autogen.gather_usage_summary() on the full agent list, including the GroupChatManager. chat_result.cost reports 0 for group chats and is not a reliable source for token counts.