# Clean Text Extraction from HTML

> Clean Text Extraction from HTML is a paid API for AI agents from api.edifiedlab.com, paid per call via x402, $0.011/call, status unknown (last checked 2026-10-02).

Strips HTML markup and returns clean, readable plain text from caller-supplied raw HTML content (no URL fetching required).

## Facts

- Endpoint: POST https://api.edifiedlab.com/v1/tools/extract-text?utm_source=zero.xyz
- Price: $0.011/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/clean-text-extraction-from-html-fed2abc6
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_wWdDR2QCiqf-1IGr4IOdH

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability clean-text-extraction-from-html-fed2abc6 -d '<json body>'
```

Example prompt: I have this raw HTML snippet and I just need the clean, human-readable text out of it — can you strip all the tags and give me the plain text?

## When to prefer this

Choose this endpoint when you already have raw HTML in memory and need to extract its readable text without fetching any URL. It is ideal for pipelines that scrape HTML externally and then need to clean it before passing to an LLM, summarizer, or search index. Prefer it over URL-based extractors when you control the fetching step yourself or when the HTML comes from a non-public source.

## Known failure modes

- HTML shorter than 24 characters returns a 400 error before any payment is charged
- Malformed or empty input rejected with 400
- Base64 decoding errors if html_b64 or content_b64 are improperly encoded
- Very large HTML payloads may hit request size limits
- Non-HTML content (e.g. plain JSON or binary) may produce garbled or empty output

## How this service works

Extract clean readable text from caller-supplied HTML (no URL fetch).

## Output

Returns clean, human-readable plain text extracted from the submitted HTML, with markup, scripts, and formatting tags removed, leaving only the readable textual content.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "html": {
   "type": "string",
   "minLength": 24,
   "description": "Decoded HTML, minimum 24 characters. Tiny fragments such as <p>hello</p> are rejected unpaid (400) before a 402 challenge."
  },
  "html_b64": {
   "type": "string",
   "description": "Base64-encoded UTF-8 HTML; decoded text must be at least 24 characters (same floor as html minLength).",
   "contentEncoding": "base64"
  },
  "content_b64": {
   "type": "string",
   "description": "Alias of html_b64; decoded text must be at least 24 characters.",
   "contentEncoding": "base64"
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/clean-text-extraction-from-html-fed2abc6/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from api.edifiedlab.com](https://www.zero.xyz/host/api.edifiedlab.com/llms.txt)
