# CRA AGENT Web Content Extractor

> CRA AGENT Web Content Extractor is a paid API for AI agents from api.cra-agent.tech, paid per call via x402, $0.005/call, status unknown (last checked 2026-10-02).

Fetches and returns clean readable text, title, description, and headings from any public URL — up to 20,000 characters, optimized for LLM context.

## Facts

- Endpoint: GET https://api.cra-agent.tech/v1/paid/web/extract?utm_source=zero.xyz
- Price: $0.005/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/cra-agent-web-content-extractor-c3418d0f
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_7RBsberLIe6JPmOM57yFp

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability cra-agent-web-content-extractor-c3418d0f
```

Example prompt: Can you fetch the readable text from https://www.arc.network and give me a clean summary of what's on the page?

## When to prefer this

Choose this endpoint when you need fast, clean text extraction from a single public web URL to feed into an LLM context. It is purpose-built for AI agent workflows, refusing private/local addresses for safety. Prefer it over generic scraping APIs when you want pre-cleaned, LLM-ready output including title, headings, and body text without writing HTML parsing logic.

## Known failure modes

- Private or local IP addresses (e.g. 192.168.x.x, localhost) are refused with an error
- Non-HTTP/HTTPS URLs or non-default ports are rejected
- Pages behind authentication or paywalls may return little or no content
- Very large or slow pages may be truncated at the 20,000 character limit
- URLs that return non-HTML content (e.g. raw PDFs, binary files) may yield minimal text

## How this service works

Title, description, headings and up to 20,000 characters of clean readable text from any public URL. Built for LLM context. Private and local addresses are refused.

## Output

Returns the page title, meta description, extracted headings, and up to 20,000 characters of clean human-readable body text from the given URL, stripped of HTML/JS/CSS — suitable for direct injection into LLM context windows.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "required": {
   "type": "string"
  },
  "properties": {
   "type": "string"
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/cra-agent-web-content-extractor-c3418d0f/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from api.cra-agent.tech](https://www.zero.xyz/host/api.cra-agent.tech/llms.txt)
