# ScrapeForAgents Web Crawler API

> ScrapeForAgents Web Crawler API is a paid API for AI agents from api.scrapeforagents.tech, paid per call via x402, $0.025/call, status unknown (last checked 2026-10-02).

Crawls a starting URL and returns structured job listing and page data up to a configurable link depth and item count.

## Facts

- Endpoint: POST https://api.scrapeforagents.tech/v1/get?utm_source=zero.xyz
- Price: $0.025/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/scrapeforagents-web-crawler-api-87f596c7
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_PxtvyvtEYKShBYIf02dJU

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability scrapeforagents-web-crawler-api-87f596c7 -d '<json body>'
```

Example prompt: Crawl get-in-it.de/jobs starting from the job search page, follow links up to 2 levels deep, collect up to 100 job listings, and give me back structured data with job titles, locations, and job IDs — stay on the same domain only.

## When to prefer this

Choose this endpoint when you need to crawl a web page or job board multiple link levels deep and receive structured, row-oriented job data (titles, locations, job IDs) without writing your own scraper. It is especially well-suited for get-in-it.de job data extraction and supports configurable crawl depth, domain restriction, and file extension filtering. Pay-per-call with no charge on failures makes it low-risk for exploratory crawls.

## Known failure modes

- Invalid or unreachable start URL returns an error or empty result set
- maxDepth set too high may cause slow or timeout responses on large sites
- Crawl blocked by site robots.txt or anti-scraping measures resulting in empty items
- Non-GET-accessible URLs (login-gated pages) return no data
- Exceeding maxItems limit truncates results without error notice

## How this service works

Pay-per-call structured web data. Failed or empty runs are not charged.

## Output

Returns a JSON object with a count of items found and an array of structured records, each containing the page URL, page name, crawl depth, job ID, job title, location, page type, and parent URL. Failed or empty crawls are not billed.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "maxDepth": {
   "type": "integer",
   "default": 2,
   "maximum": 10,
   "minimum": 0,
   "description": "How many link levels to follow after the start URL. Zero returns only the start URL."
  },
  "maxItems": {
   "type": "integer",
   "default": 100,
   "minimum": 0,
   "description": "Maximum output rows across the crawl. Zero removes the cap."
  },
  "startUrl": {
   "type": "string",
   "description": "Public GET in IT URL to crawl. Use the job search for current vacancies or sitemap.xml for the public URL inventory."
  },
  "sameDomainOnly": {
   "type": "boolean",
   "default": true,
   "description": "Follow only links on get-in-it.de and its www host."
  },
  "allowDuplicates": {
   "type": "boolean",
   "default": false,
   "description": "Include a URL again when it is linked from another parent; each URL is still fetched at most once."
  },
  "ignoredExtensions": {
   "type": "array",
   "default": [
    "pdf",
    "jpg",
    "jpeg",
    "png",
    "gif",
    "svg",
    "webp",
    "zip",
    "css",
    "js",
    "woff",
    "woff2"
   ],
   "description": "Skip links whose paths end in these file extensions. Enter names without the dot."
  },
  "maxChildrenPerLink": {
   "type": "integer",
   "default": 50,
   "maximum": 10000,
   "minimum": 1,
   "description": "Maximum distinct links taken from each page, including job search results."
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "count": 1,
  "items": [
   {
    "url": null,
    "name": null,
    "depth": null,
    "jobId": null,
    "jobTitle": null,
    "location": null,
    "pageType": null,
    "parentUrl": null
   }
  ],
  "product": "get"
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/scrapeforagents-web-crawler-api-87f596c7/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from api.scrapeforagents.tech](https://www.zero.xyz/host/api.scrapeforagents.tech/llms.txt)
