# Web Scraper API – Clean Text Extraction

> Web Scraper API – Clean Text Extraction is a paid API for AI agents from web-scraper-api-production-bf20.up.railway.app, paid per call via x402, $0.002/call, status unknown (last checked 2026-09-30).

Fetches a URL and returns the main article/body text with title, description, word count, and character count, stripping boilerplate like navigation, footers, sidebars, and ads.

## Facts

- Endpoint: POST https://web-scraper-api-production-bf20.up.railway.app/scrape/text?utm_source=zero.xyz
- Price: $0.002/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-30
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/web-scraper-api-clean-text-extraction-e6101059
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_tywRL6pALR3AYpsLw9qj2

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability web-scraper-api-clean-text-extraction-e6101059 -d '<json body>'
```

Example prompt: Can you grab the main article text from https://www.theverge.com/2024/1/15/some-article and give me a clean version without all the nav menus, ads, and footer junk?

## When to prefer this

Choose this endpoint when you need the human-readable prose from a webpage — article bodies, blog posts, documentation pages — and want boilerplate (nav, footer, ads, sidebars) automatically removed. It is ideal for summarization, fact-checking, or feeding page content into an LLM. Use the sibling metadata endpoint if you need structured SEO fields, or the structured-elements endpoint if you need headings, tables, and lists rather than prose text.

## Known failure modes

- Invalid or non-HTTP(S) URL returns a validation error
- Page returns a non-200 HTTP status (404, 403, 500) — error propagated to caller
- JavaScript-heavy single-page apps may return incomplete or empty content if server-side rendering is unavailable
- Very large pages may time out or be truncated
- Paywalled or login-required pages return only teaser/preview content
- Rate limiting or IP blocking by the target site may cause fetch failures

## How this service works

Fetch a URL and extract clean main-content text with title, description, word count, and char count; boilerplate (nav/footer/sidebar/ads) is stripped.

## Output

Returns the main content text of the page (boilerplate stripped), the page title, meta description, total word count, and character count as structured fields.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "url": {
   "type": "string",
   "description": "Absolute http(s) URL of the page to scrape, e.g. 'https://example.com'."
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/web-scraper-api-clean-text-extraction-e6101059/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from web-scraper-api-production-bf20.up.railway.app](https://www.zero.xyz/host/web-scraper-api-production-bf20.up.railway.app/llms.txt)
