# Tollkit Extract /read – Web Page Text Extractor

> Tollkit Extract /read – Web Page Text Extractor is a paid API for AI agents from extract.tollkit.dev, paid per call via x402, $0.005/call, status unknown (last checked 2026-10-02).

Fetches any public web page (including JS-rendered pages), strips boilerplate, and returns clean visible text, page title, and final URL — up to 50,000 characters.

## Facts

- Endpoint: POST https://extract.tollkit.dev/read?utm_source=zero.xyz
- Price: $0.005/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/tollkit-extract-read-web-page-text-extractor-713929db
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_Pcc6u_4C1dLCfJvKMkEF_

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability tollkit-extract-read-web-page-text-extractor-713929db -d '<json body>'
```

Example prompt: Can you read the full text of this page for me — https://www.theverge.com/2024/5/1/example-article — and give me the clean article content without all the navigation and ads?

## When to prefer this

Choose this endpoint when you need clean, readable text from any public web page, especially JavaScript-rendered pages where a simple HTTP GET would return incomplete content. Ideal for RAG pipelines, LLM document ingestion, research workflows, or page monitoring. Prefer over generic HTTP fetch tools when you need boilerplate-stripped prose rather than raw HTML. Use the companion /summarize endpoint if you need a condensed version rather than the full text.

## Known failure modes

- Page fails to load or is unreachable — not charged, returns error
- Page is behind a login/paywall — may return empty or partial content
- Page exceeds 50,000 character limit — content is truncated
- Invalid or non-http(s) URL — returns validation error
- JavaScript-heavy SPA that never settles — may timeout

## How this service works

Read any web page as clean text: give a URL, get the page's visible text, title and final URL. Renders JavaScript pages in a real browser first, strips navigation and boilerplate, up to 50,000 characters. Use for web reading, research, RAG, page monitoring and feeding a page to an LLM (pairs with /summarize). $0.005 per page; a page that fails to load, errors or is empty is not charged.

## Output

Returns the page's clean visible text (up to 50,000 characters), the page title, and the final URL after any redirects. JavaScript is executed before extraction to capture dynamically rendered content. Navigation menus, footers, and other boilerplate are stripped. Failed, empty, or erroring pages are not charged.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "url": {
   "type": "string",
   "format": "uri",
   "description": "Absolute http(s) URL of a public page. JavaScript pages are rendered first."
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/tollkit-extract-read-web-page-text-extractor-713929db/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from extract.tollkit.dev](https://www.zero.xyz/host/extract.tollkit.dev/llms.txt)
