# x402.orthogonal.com Web Crawl & Structured Data Extractor

> x402.orthogonal.com Web Crawl & Structured Data Extractor is a paid API for AI agents from x402.orthogonal.com, paid per call via x402, $0.03/call, status unknown (last checked 2026-10-02).

Crawls a website starting from a given URL, intelligently prioritizes relevant internal links using a JSON Schema and instructions, and extracts structured data from selected pages.

## Facts

- Endpoint: POST https://x402.orthogonal.com/context-dev/web/extract?utm_source=zero.xyz
- Price: $0.03/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/x402-orthogonal-com-web-crawl-structured-data-extractor-6279675d
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap__UWvVe4tY7AGJw71wExw6

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability x402-orthogonal-com-web-crawl-structured-data-extractor-6279675d -d '<json body>'
```

Example prompt: Crawl https://acme.com/products and extract a list of all products with their name, price, and description — look up to 3 levels deep, check up to 20 pages, and use the schema {products: [{name, price, description}]}; make sure every value is grounded in what's actually on the page.

## When to prefer this

Choose this endpoint when you need to extract structured, schema-conformant data from an entire website or a multi-page section of one — not just a single URL. It is especially valuable when you need intelligent link prioritization (using instructions and a schema to guide which pages to follow), fact-grounded extraction, and control over crawl depth and page budget. Prefer it over a single-page scraper when the information is spread across multiple pages or when you want a structured JSON object rather than raw HTML or Markdown.

## Known failure modes

- URL is unreachable or returns non-200 status — crawl fails with no data
- Schema is too complex or ambiguous — extraction may return partial or empty fields
- maxPages cap of 50 hit before full coverage — some data may be missing
- stopAfterMs budget exhausted — crawl stops early, returning partial results
- Page requires JavaScript-heavy interaction beyond waitForMs — data not rendered
- Site blocks crawlers via robots.txt or rate limiting — incomplete or failed extraction
- factCheck enabled but values not found on page — fields returned as null or omitted

## How this service works

Crawl a website, use the provided JSON Schema and instructions to prioritize relevant internal links, and extract structured data from the selected pages.

## Output

A JSON object conforming to the provided schema, populated with data extracted from the crawled pages. Values are grounded in content found on the site when factCheck is enabled. The response contains the structured fields as specified in the input schema.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "url": {
   "type": "string",
   "description": "The starting website URL to crawl and extract from. Must include http:// or https://."
  },
  "schema": {
   "type": "object",
   "description": "JSON Schema for the returned data object."
  },
  "maxAgeMs": {
   "type": "integer",
   "description": "Return cached scrape results if younger than this many milliseconds."
  },
  "maxDepth": {
   "type": "integer",
   "description": "Optional maximum link depth from the starting URL (0 = only the starting page)."
  },
  "maxPages": {
   "type": "integer",
   "description": "Maximum number of pages to analyze for extraction. Hard cap: 50. Defaults to 5."
  },
  "factCheck": {
   "type": "boolean",
   "description": "When true, every returned value must be grounded in facts stated on the page."
  },
  "waitForMs": {
   "type": "integer",
   "description": "Optional browser wait time in milliseconds after initial page load for each crawled page."
  },
  "stopAfterMs": {
   "type": "integer",
   "description": "Soft time budget for the crawl in milliseconds. Min: 10000. Max: 110000. Default: 80000."
  },
  "instructions": {
   "type": "string",
   "description": "Optional extraction guidance, such as which facts to prioritize or how to interpret fields."
  },
  "includeFrames": {
   "type": "boolean",
   "description": "When true, iframe contents are included in Markdown before extraction."
  },
  "followSubdomains": {
   "type": "boolean",
   "description": "When true, follow links on subdomains of the starting URL's domain."
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/x402-orthogonal-com-web-crawl-structured-data-extractor-6279675d/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from x402.orthogonal.com](https://www.zero.xyz/host/x402.orthogonal.com/llms.txt)
