# readpage Web Content Extractor

> readpage Web Content Extractor is a paid API for AI agents from readpage.x402utils.workers.dev, paid per call via x402, $0.001/call, status unknown (last checked 2026-10-02).

Fetches a public web page by URL and returns clean, structured text including title, description, headings, main content, and links — with scripts, navigation, and footers stripped.

## Facts

- Endpoint: GET https://readpage.x402utils.workers.dev/v1/meta?utm_source=zero.xyz
- Price: $0.001/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/readpage-web-content-extractor-d654296d
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_sP32evwmaMk1xldDjEMTr

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability readpage-web-content-extractor-d654296d
```

Example prompt: Can you fetch the content of https://example.com/article and give me the title, main text, and any important links — with all the navigation and scripts stripped out?

## When to prefer this

Choose this endpoint when an agent needs to read and understand the textual content of a specific public URL without dealing with raw HTML, JavaScript, or boilerplate — especially for article extraction, research, or link harvesting tasks at $0.001 per call with no setup required.

## Known failure modes

- Non-public or login-gated URLs return empty or partial content
- Invalid or malformed URLs cause a request failure
- Pages with heavy JavaScript rendering may not yield full content (server-side rendering required)
- Rate limiting or bot-blocking by target site can result in no content
- HTTP 402 payment required if USDC micropayment via x402 fails

## How this service works

Turn any public web page into clean, agent-ready text: title, description, headings, main text and links. Scripts, navigation and footers are stripped. Paid per call via x402 (USDC on Base).

## Output

A JSON object containing the page title, meta description, OpenGraph tags, main body text, headings, and links extracted from the target URL — with boilerplate like scripts, navigation, and footers removed.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "required": [
  "input"
 ],
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "type",
    "method"
   ],
   "properties": {
    "type": {
     "type": "string",
     "const": "http"
    },
    "method": {
     "enum": [
      "GET",
      "HEAD",
      "DELETE"
     ],
     "type": "string"
    },
    "queryParams": {
     "type": "object",
     "required": [
      "url"
     ],
     "properties": {
      "url": {
       "type": "string",
       "description": "http(s) URL"
      }
     }
    }
   },
   "additionalProperties": false
  },
  "output": {
   "type": "object",
   "required": [
    "type"
   ],
   "properties": {
    "type": {
     "type": "string"
    },
    "example": {
     "type": "object"
    }
   }
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "og": {},
  "title": "Example Domain",
  "description": null
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/readpage-web-content-extractor-d654296d/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from readpage.x402utils.workers.dev](https://www.zero.xyz/host/readpage.x402utils.workers.dev/llms.txt)
