# Readable Text Extractor

> Readable Text Extractor is a paid API for AI agents from scrooge-x402-tool-mill.vercel.app, paid per call via x402, $0.05/call, status unknown (last checked 2026-10-01).

Fetches a public HTML page and returns its main readable plaintext content along with title, language, word count, and fetch timestamp

## Facts

- Endpoint: GET https://scrooge-x402-tool-mill.vercel.app/v1/readable-text?utm_source=zero.xyz
- Price: $0.05/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-01
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/readable-text-extractor-93a6735c
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_Nks0ZNWGcuQFuCaHkmTJx

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability readable-text-extractor-93a6735c
```

Example prompt: Can you pull the readable text from https://www.bbc.com/news/science-environment-12345678 — limit it to 8000 characters so I can summarize it?

## When to prefer this

Choose this endpoint when you need clean, human-readable plaintext from a public webpage — ideal for feeding article text into LLMs, summarization pipelines, or NLP tasks. It strips away HTML boilerplate and returns only the main content, making it more useful than raw HTML fetchers when you need readable prose. Prefer it over link-graph or metadata endpoints when your goal is the body text, not the structure or metadata of the page. It requires no API key and charges a flat $0.05 USDC per call on Base.

## Known failure modes

- Private or link-local URLs (e.g. localhost, 192.168.x.x) are blocked with an error — only public http/https URLs accepted
- Pages that require JavaScript rendering may return incomplete or empty text since the fetcher processes static HTML
- Very large pages may return truncated text if the readable content exceeds the maxChars cap
- Non-HTML responses (PDFs, images, APIs) may yield empty or garbled text output
- Payment failure or insufficient USDC balance on Base will prevent the request from completing
- Paywalled or login-gated pages may return minimal or no readable content

## How this service works

Extract the main readable plaintext of a public HTML page: final url, title, text (capped), wordCount, language, and fetchedAt. POST JSON {"url":"https://example.com","maxChars":12000}. Optional maxChars is an integer 1–12000 (default 12000). Public http(s) only; private hosts blocked. $0.05 USDC on Base (eip155:8453). No API key.

## Output

Returns a JSON object with the final URL after redirects, the extracted plaintext body (capped at maxChars, default 12000), the HTML page title (or null), detected language code (or null), word count as an integer, and an ISO-8601 timestamp of when the page was fetched.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "required": [
  "input"
 ],
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "type",
    "method"
   ],
   "properties": {
    "type": {
     "type": "string",
     "const": "http"
    },
    "method": {
     "enum": [
      "GET"
     ],
     "type": "string"
    },
    "queryParams": {
     "type": "object",
     "properties": {
      "url": {
       "type": "string",
       "format": "uri",
       "description": "Public http(s) URL to fetch (SSRF-safe; private/link-local blocked)"
      },
      "maxChars": {
       "type": "string",
       "description": "Optional maxChars query string (1–12000; handler clamps)"
      }
     }
    }
   },
   "additionalProperties": false
  },
  "output": {
   "type": "object",
   "required": [
    "type"
   ],
   "properties": {
    "type": {
     "type": "string"
    },
    "example": {
     "type": "object",
     "required": [
      "url",
      "title",
      "text",
      "wordCount",
      "language",
      "fetchedAt"
     ],
     "properties": {
      "url": {
       "type": "string",
       "format": "uri",
       "description": "Final URL after redirects"
      },
      "text": {
       "type": "string",
       "description": "Main readable plaintext, capped by maxChars (default and max 12000)"
      },
      "title": {
       "type": [
        "string",
        "null"
       ],
       "description": "HTML title, or null"
      },
      "language": {
       "type": [
        "string",
        "null"
       ],
       "description": "html lang or content-language, or null"
      },
      "fetchedAt": {
       "type": "string",
       "format": "date-time",
       "description": "ISO-8601 time the page was fetched"
      },
      "wordCount": {
       "type": "integer",
       "minimum": 0
      }
     },
     "additionalProperties": false
    }
   }
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/readable-text-extractor-93a6735c/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from scrooge-x402-tool-mill.vercel.app](https://www.zero.xyz/host/scrooge-x402-tool-mill.vercel.app/llms.txt)
