# Introducing Batch: send a list, collect the results

> Fetching 100,000 URLs is not a bigger version of fetching one. Batch is the queue, the retries and the per-URL accounting, run by us, so you submit a list and pick the results up when it finishes.

Source: https://www.zenrows.com/blog/introducing-batch

Fetching a hundred thousand URLs is not a bigger version of fetching one. It needs a queue. It needs retries that do not re-run what already worked. It needs somewhere to put the results, back-pressure when a target slows down, and some way to find out the run finished without holding a connection open for hours.

Every team that reaches that scale ends up building the same layer, around a synchronous API that was never designed for it. The scraper is rarely the hard part. The machinery around it is.

**Today we're launching Batch. Submit your URLs as one job, walk away, and collect the results when it finishes.**

## What is Batch?

Batch is the orchestration layer over [Fetch](/products/fetch). You hand it a set of URLs as one job instead of holding a request open per URL.

Three nouns describe the whole model:

- A **Job** carries your URLs and their scraping options.
- A **Run** is one execution of that job.
- A **Task** is one URL, with its own final status, real HTTP code, and either a result or a structured error.

Submit up to 100,000 URLs in a single job. It runs on the same API key as Fetch, so there is nothing new to set up:

```python
import requests

API_KEY = "YOUR_API_KEY"

resp = requests.post(
    "https://async.api.zenrows.com/v1/jobs",
    headers={"X-API-Key": API_KEY},
    json={
        "tasks": [
            {"url": "https://example.com/product/1", "external_id": "sku-1"},
            {"url": "https://example.com/product/2", "external_id": "sku-2"},
        ]
    },
)

job_id = resp.json()["job_id"]
print(job_id)
```

The call returns as soon as the work is accepted. It does not wait for the scraping to finish, which is the whole point.

## How Batch works

### Every URL ends with an answer

Each task finishes `successful` or `failed`, carries the real HTTP status code, and a failed task carries a structured error with a code and a detail. There is no silent shortfall to reconcile at the end, and no guessing which of your 100,000 URLs never came back.

Pass an `external_id` on any task and it comes back verbatim on the result row, so joining results to your own records is a dictionary lookup rather than URL string matching:

```python
results = requests.get(
    f"https://async.api.zenrows.com/v1/jobs/{job_id}/results",
    headers={"X-API-Key": API_KEY},
).json()

for row in results["results"]:
    print(row["external_id"], row["status"], row["result_url"])
```

Each `result_url` is a presigned link to that page's content. Re-list the results whenever you need a fresh one.

### Reruns only touch what failed

When a run finishes with failures, rerun the failed tasks alone. The successful results carry over untouched, and you are not paying to scrape 97,000 pages a second time to recover 3,000. Transient failures are retried inside the run before they ever reach you.

### One job, many domains

A job is a list of URLs plus options, so the targets can be mixed. Put twelve different sites in one job and override the scraping parameters on the individual URLs that need different handling. You do not need one pipeline per site shape.

### You know the cost before you submit

Estimate a job's cost client side before you send it, then read the actual spend on the run afterwards. Batch adds no charge of its own: you pay for the Fetch calls it runs and nothing on top, and only successful requests are charged.

## What you can build with Batch

**Refresh a whole catalog.** Point one job at every product URL you track, run it, and read per-URL status to see exactly which SKUs came back and which need another pass.

**Seed a RAG index.** Send the whole corpus as one job, collect clean page content when it finishes, and use the per-task accounting to prove your index is complete rather than assuming it.

**Retire your queue.** If you already run Celery or a thread pool around a scraping API, Batch is that component, managed. The dispatcher, the retry logic and the result store stop being code you maintain.

**Enrich a list you were handed.** Drop in a list of company URLs, walk away, and come back to structured results with the failures already labelled.

## Try it today

Batch is live for every Zenrows account with an API key.

Send your first job with the snippet above, or read the guide for the full API.

[Start building free](https://app.zenrows.com/register) · [Read the docs](https://docs.zenrows.com/batch/introduction) · [Batch product page](/products/batch)
