# Web Page to Clean Text Extractor

> Web Page to Clean Text Extractor is a paid API for AI agents from twin.unykorn.org, paid per call via x402, $0.004/call, status unknown (last checked 2026-10-01).

Fetches any public URL and returns cleaned, readable text including title, description, headings, links, and body content

## Facts

- Endpoint: POST https://twin.unykorn.org/web-extract?utm_source=zero.xyz
- Price: $0.004/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-01
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/web-page-to-clean-text-extractor-c425b6d0
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_BvMsfmjdmWlZko5o0YL5m

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability web-page-to-clean-text-extractor-c425b6d0 -d '<json body>'
```

Example prompt: Can you pull the clean readable text from https://www.bbc.com/news/technology-12345678 — include the title, description, headings, and links, up to 10000 characters?

## When to prefer this

Choose this endpoint when you need to quickly extract readable human-facing text from any public URL without running a full browser — ideal for LLM context feeding, article summarization, link extraction, or competitive research where you need clean text rather than raw HTML. It is especially useful in agent pipelines that need to process web content programmatically at low cost ($0.004 per call).

## Known failure modes

- URL is not publicly accessible (paywalled, behind login, or blocked) — returns error or empty content
- URL is malformed or invalid — returns validation error
- Page returns non-HTML content (PDF, binary) — may return empty or partial text
- max_chars too low to capture meaningful content — truncated output
- Server timeout if target page is slow to respond
- Bot-blocking or rate-limiting by the target website — returns error or incomplete content

## How this service works

Web page to clean text: title, description, headings, links and readable text from any public URL — Genesis402 / UnyKorn Operator Network

## Output

Returns the page's title, meta description, structured headings, main readable body text, and optionally all hyperlinks found on the page — all cleaned of HTML markup and limited to the requested character count.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "params": {
   "type": "object",
   "properties": {
    "url": {
     "type": "string",
     "format": "uri"
    },
    "max_chars": {
     "type": "integer",
     "maximum": 60000,
     "minimum": 500
    },
    "include_links": {
     "type": "boolean"
    }
   }
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "ok": true,
  "text": "<readable text>",
  "type": "web-extract",
  "links": [
   {
    "url": "https://…",
    "text": "<anchor>"
   }
  ],
  "title": "<title>",
  "receipt": {
   "tx_hash": "0x<64hex>",
   "amount_usd": 0.004,
   "receipt_id": "g402-<16hex>"
  },
  "headings": [
   {
    "text": "<h1>",
    "level": 1
   }
  ],
  "final_url": "https://www.x402.org/",
  "description": "<meta description>",
  "http_status": 200,
  "content_sha256": "<64hex>"
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/web-page-to-clean-text-extractor-c425b6d0/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from twin.unykorn.org](https://www.zero.xyz/host/twin.unykorn.org/llms.txt)
