# Web Page Extractor — Clean Markdown via URL Fetch

> Web Page Extractor — Clean Markdown via URL Fetch is a paid API for AI agents from toolbelt402.tpoborne.workers.dev, paid per call via x402, $0.02/call, status unknown (last checked 2026-10-01).

Fetches up to 5 URLs and returns the main article content as clean Markdown with title, author, date, links, images, word count, and reading time — stripping nav, ads, and footers.

## Facts

- Endpoint: POST https://toolbelt402.tpoborne.workers.dev/web/extract?utm_source=zero.xyz
- Price: $0.02/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-01
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/web-page-extractor-clean-markdown-via-url-fetch-97b92290
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_iAmCL_xw_P1x9gN3wMt25

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability web-page-extractor-clean-markdown-via-url-fetch-97b92290 -d '<json body>'
```

Example prompt: Can you read this article for me and give me the main content as clean text? Here's the URL: https://example.com/some-long-article — strip out the ads, navigation, and footer noise.

## When to prefer this

Use this endpoint when an agent needs to read and use the textual content of one or more web pages without consuming excessive context tokens. It is ideal over raw HTTP fetching because it automatically strips boilerplate (nav, ads, footers) and returns structured Markdown plus metadata. Prefer it over browser-based scraping tools when JavaScript rendering is not required and robots.txt compliance is acceptable. Best for article reading, research pipelines, and content summarization workflows.

## Known failure modes

- URL is blocked by robots.txt — endpoint honors robots.txt by default and will refuse or return an error
- URL is unreachable or returns a non-200 HTTP status
- Page requires JavaScript rendering to load content — server-side fetch may return empty or incomplete content
- Paywall or login-gated content returns minimal extractable text
- More than 5 URLs submitted — batch limit exceeded
- Malformed URL input causes a 400-level validation error

## How this service works

Read a web page for me: fetches up to 5 URLs and returns the main article content as clean Markdown (nav, ads, and footers stripped) plus title, author, published date, canonical URL, links, images, word count and reading time. Typically 15-30x smaller than the raw HTML, so it saves far more in context tokens than it costs. Honors robots.txt by default. No browser or network access needed on your side.

## Output

A structured response per URL containing: cleaned Markdown body of the main article content, title, author name, published date, canonical URL, list of links and images found, word count, and estimated reading time. Nav bars, ads, and footers are removed. Content is typically 15-30x smaller than raw HTML.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "url": {
   "type": "string",
   "description": "Single URL to extract (or use urls)"
  },
  "urls": {
   "type": "array",
   "items": {
    "type": "string"
   },
   "maxItems": 5,
   "description": "Up to 5 absolute http(s) URLs"
  },
  "format": {
   "enum": [
    "markdown",
    "text"
   ],
   "type": "string",
   "description": "Output format (default markdown)"
  },
  "maxBytes": {
   "type": "integer",
   "description": "Cap HTML read per page (default 524288)"
  },
  "userAgent": {
   "type": "string",
   "description": "User-agent to send and to evaluate robots.txt against"
  },
  "includeLinks": {
   "type": "boolean",
   "description": "Include links found in the article body (default true)"
  },
  "includeImages": {
   "type": "boolean",
   "description": "Include images found in the article body (default true)"
  },
  "respectRobots": {
   "type": "boolean",
   "description": "Honor robots.txt for the given user-agent (default true)"
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/web-page-extractor-clean-markdown-via-url-fetch-97b92290/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from toolbelt402.tpoborne.workers.dev](https://www.zero.xyz/host/toolbelt402.tpoborne.workers.dev/llms.txt)
