# APEX Faucet Site Extract API

> APEX Faucet Site Extract API is a paid API for AI agents from apexfaucet.xyz, paid per call via x402, $0.14/call, status unknown (last checked 2026-10-03).

Crawls and extracts text content from pages on a given host, returning structured page data including titles, URLs, and character counts

## Facts

- Endpoint: GET https://apexfaucet.xyz/api/x402/site-extract?utm_source=zero.xyz
- Price: $0.14/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-03
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/apex-faucet-site-extract-api-30222e47
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_npIaBGqHqxJzD2VmyfM7K

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability apex-faucet-site-extract-api-30222e47
```

Example prompt: Can you crawl docs.arc.io for me and extract the text content from up to 10 pages, starting from the homepage?

## When to prefer this

Use this endpoint when you need to extract multi-page text content from a website in a structured format, especially when you want host-scoped crawling with robots.txt compliance. Prefer this over generic scraping tools when you need page-level metadata (title, URL, character count) and want to respect crawling rules automatically. It is particularly useful for indexing documentation sites or gathering site content for downstream analysis by an AI agent.

## Known failure modes

- URL is not reachable or returns non-200 status
- robots.txt disallows crawling, resulting in zero pages fetched
- Host mismatch — only pages on the same host as the start URL are fetched
- Page count exceeds hard cap of 25, capped silently
- Payment of 0.14 USDC not completed, resulting in 402 error
- Invalid or malformed URL input causes error response

## How this service works

Website crawler: up to 25 pages of one site as clean text, robots.txt obeyed. Render a whole section of a site - up to 25 pages - in a real browser and return every page as clean text.

## Output

A JSON object with ok status, and a data object containing the host, start URL, an array of pages (each with URL, title, and character count), a list of failed URLs, robots.txt metadata, total page count, and total character count across all fetched pages.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "required": [
  "input"
 ],
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "type",
    "method",
    "queryParams"
   ],
   "properties": {
    "type": {
     "type": "string",
     "const": "http"
    },
    "method": {
     "enum": [
      "GET"
     ],
     "type": "string"
    },
    "queryParams": {
     "type": "object",
     "required": [
      "url"
     ],
     "properties": {
      "url": {
       "type": "string",
       "description": "The page to start from. Only pages on this same host are fetched, and robots.txt is obeyed."
      },
      "full": {
       "type": "string",
       "description": "Set to 1 to render every planned page ($0.14 over x402). Programs pay for every call, with or without it; only a browser sees the crawl plan without it."
      },
      "pages": {
       "type": "string",
       "pattern": "^[0-9]{1,2}$",
       "description": "How many pages to render, default 10, hard cap 25 (a query string, so digits as text)."
      }
     }
    }
   }
  },
  "output": {
   "type": "object"
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "schema": {
  "type": "object",
  "properties": {
   "ok": {
    "type": "boolean"
   },
   "data": {
    "type": "object",
    "properties": {
     "host": {
      "type": "string"
     },
     "pages": {
      "type": "array"
     },
     "start": {
      "type": "string"
     },
     "failed": {
      "type": "array"
     },
     "robots": {
      "type": "object"
     },
     "pageCount": {
      "type": "integer"
     },
     "totalCharacters": {
      "type": "integer"
     }
    }
   },
   "paid": {
    "type": "object"
   }
  }
 },
 "example": {
  "ok": true,
  "data": {
   "host": "docs.arc.io",
   "pages": [
    {
     "url": "https://docs.arc.io/",
     "title": "Welcome to Arc docs",
     "characters": 1384
    }
   ],
   "start": "https://docs.arc.io/",
   "pageCount": 5,
   "totalCharacters": 28012
  },
  "paid": {
   "asset": "USDC",
   "amount": 1,
   "network": "eip155:5042"
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/apex-faucet-site-extract-api-30222e47/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from apexfaucet.xyz](https://www.zero.xyz/host/apexfaucet.xyz/llms.txt)
