# x402 AgentUtility Webpage Scraper

> x402 AgentUtility Webpage Scraper is a paid API for AI agents from x402.agentutility.ai, paid per call via x402, $0.04/call, status unknown (last checked 2026-10-02).

Scrapes a single webpage and returns its title, metadata, headings, body content (text/HTML/markdown), and outbound links using Cheerio-based server-side rendering.

## Facts

- Endpoint: POST https://x402.agentutility.ai/scrape?utm_source=zero.xyz
- Price: $0.04/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/x402-agentutility-webpage-scraper-b8ca3f73
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_EdKsdwS2u9UpO_SdGy4bM

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability x402-agentutility-webpage-scraper-b8ca3f73 -d '<json body>'
```

Example prompt: Can you scrape https://techcrunch.com/2024/01/15/openai-news/ and give me the page title, description, main headings, and body text as clean markdown?

## When to prefer this

Choose this endpoint for static pages, server-side rendered sites, and any URL where content is available in the initial HTML response. It is significantly faster and cheaper ($0.04 USDC) than headless-browser alternatives. Avoid it for JavaScript-heavy SPAs where content is rendered client-side — use a browser-based screenshot or rendering service instead.

## Known failure modes

- URL is unreachable or returns non-200 HTTP status — endpoint returns an error with the HTTP status code
- Page is a JavaScript-heavy SPA that requires a real browser — content may be empty or minimal since no headless browser is used
- Malformed or invalid URL input — returns validation error
- Page blocks server-side scraping via robots.txt or IP blocking — may return empty or error response
- Payment failure (x402) — request is rejected before scraping begins

## How this service works

Scrape any webpage. Pulls title, description, canonical URL, OpenGraph + Twitter card metadata, headings, and outbound links from a single URL. Server-side rendering. Body content rendered as text / raw HTML / clean markdown. Optional link extraction. Cheerio-based, no headless browser — fast and cheap, ideal for static pages and SSR sites. Alias of scrape-website. For JS-heavy SPAs that need a real browser, see website-screenshot.

## Output

Returns structured data including page title, meta description, canonical URL, OpenGraph and Twitter card metadata fields, all heading tags (H1–H6), body content in the requested format (plain text, raw HTML, or clean markdown), and optionally a list of outbound hyperlinks found on the page.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "required": [
  "input"
 ],
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "type",
    "method",
    "bodyType",
    "body"
   ],
   "properties": {
    "body": {
     "required": [
      "url"
     ],
     "properties": {
      "url": {
       "type": "string",
       "description": "Public URL to fetch and parse. Must include scheme (http/https). Follows redirects."
      },
      "format": {
       "enum": [
        "text",
        "html",
        "markdown"
       ],
       "type": "string",
       "description": "Body output format. 'text' (default), 'html' (raw), or 'markdown' (clean — best for LLM ingestion)."
      },
      "user_agent": {
       "type": "string",
       "description": "Custom User-Agent header. Defaults to a modern desktop Chrome UA."
      },
      "include_links": {
       "type": "boolean",
       "description": "If true, also returns an array of all <a href> links on the page. Default false."
      }
     }
    },
    "type": {
     "type": "string",
     "const": "http"
    },
    "method": {
     "enum": [
      "POST"
     ],
     "type": "string"
    },
    "bodyType": {
     "enum": [
      "json",
      "form-data",
      "text"
     ],
     "type": "string"
    }
   },
   "additionalProperties": false
  },
  "output": {
   "type": "object",
   "required": [
    "type"
   ],
   "properties": {
    "type": {
     "type": "string"
    },
    "example": {
     "type": "object",
     "properties": {
      "h1": {
       "type": "string"
      },
      "og": {
       "type": "object",
       "properties": {}
      },
      "url": {
       "type": "string"
      },
      "lang": {
       "type": "string"
      },
      "text": {
       "type": "string"
      },
      "title": {
       "type": "string"
      },
      "format": {
       "type": "string"
      },
      "twitter": {
       "type": "object",
       "properties": {}
      },
      "canonical": {
       "type": "null"
      },
      "final_url": {
       "type": "string"
      },
      "body_chars": {
       "type": "integer"
      },
      "description": {
       "type": "string"
      },
      "status_code": {
       "type": "integer"
      }
     }
    }
   }
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "h1": "Example Domain",
  "og": {},
  "url": "https://example.com",
  "lang": "en",
  "text": "Example Domain\n\nThis domain is for use in documentation examples without needing permission. Avoid use in operations.\n\nLearn more",
  "title": "Example Domain",
  "format": "text",
  "twitter": {},
  "canonical": null,
  "final_url": "https://example.com/",
  "body_chars": 128,
  "description": "",
  "status_code": 200
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/x402-agentutility-webpage-scraper-b8ca3f73/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from x402.agentutility.ai](https://www.zero.xyz/host/x402.agentutility.ai/llms.txt)
