# HorizonPulse HTML Extractor

> HorizonPulse HTML Extractor is a paid API for AI agents from horizonpulse.dev, paid per call via x402, $0.015/call, status unknown (last checked 2026-10-01).

Extracts structured fields (title, description, canonical, links, images, headings, JSON-LD, text sample) from a public URL or raw HTML, with SSRF protection

## Facts

- Endpoint: GET https://horizonpulse.dev/api/extract?utm_source=zero.xyz
- Price: $0.015/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-01
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/horizonpulse-html-extractor-5aa58644
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_h4Vv2XLzxAS7vEBBU9zoN

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability horizonpulse-html-extractor-5aa58644
```

Example prompt: Can you extract all the structured metadata from https://example.com/article — I need the title, description, canonical URL, headings, images, links, and any JSON-LD data on the page.

## When to prefer this

Choose this endpoint when you need multiple structured fields from a webpage in one call (title, description, headings, links, images, JSON-LD) rather than just clean text. It is preferable to the sibling markdown/text-fetch endpoint when you need metadata, link lists, or structured schema.org data. It is ideal for SEO analysis, content enrichment, link graph construction, or validating page metadata. Use it over a generic HTTP proxy when you specifically need parsed, structured HTML fields rather than raw response bodies.

## Known failure modes

- Private or localhost URLs are blocked and return an SSRF-safety error
- Invalid or malformed URLs result in a fetch error
- HTML size exceeding the cap is truncated or rejected
- Pages behind authentication or paywalls may return incomplete or empty content
- Non-HTML content types (PDF, images) may yield minimal extraction results
- Network timeouts on slow or unreachable public URLs

## How this service works

Pay-per-call APIs for AI agents via x402 on Base: web fetch, HTTP proxy, page extract, and crypto market data.

## Output

A structured JSON object containing extracted page fields: title, meta description, canonical URL, all links, image URLs, heading hierarchy, JSON-LD structured data blocks, and a plain-text sample of the page body. Private/localhost URLs are blocked for SSRF safety.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "required": [
  "input"
 ],
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "type",
    "method"
   ],
   "properties": {
    "type": {
     "type": "string",
     "const": "http"
    },
    "method": {
     "enum": [
      "GET",
      "HEAD",
      "DELETE"
     ],
     "type": "string"
    },
    "queryParams": {
     "type": "object",
     "required": [
      "url"
     ],
     "properties": {
      "url": {
       "type": "string",
       "description": "Absolute http(s) URL to fetch and extract (required on GET; on POST send url or html in the JSON body). Private/localhost blocked."
      },
      "html": {
       "type": "string",
       "description": "POST /api/extract JSON body only (ignored on GET): raw HTML to parse instead of fetching (size-capped). If both url and html are sent, html is parsed and url is echoed."
      },
      "fields": {
       "type": "object",
       "description": "Optional CSS-selector fields (max 20). Map of name to selector string or {selector, attr?: \"text\"|\"html\"|<attribute>, all?: boolean, limit?: 1-50}. GET: URL-encoded JSON. Missing fields come back null with fieldErrors; if none match, 422 no_fields_matched and no charge."
      }
     }
    }
   },
   "additionalProperties": false
  },
  "output": {
   "type": "object",
   "required": [
    "type"
   ],
   "properties": {
    "type": {
     "type": "string"
    },
    "example": {
     "type": "object",
     "description": "Horizon Pulse structured HTML extract result. The example is a trimmed real response recorded from GET /api/demo/extract at 2026-10-01T14:36Z UTC; live values differ on every call."
    }
   }
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "ok": true,
  "url": "https://horizonpulse.dev/",
  "links": [
   {
    "href": "https://horizonpulse.dev/",
    "text": "Horizon Pulse"
   },
   {
    "href": "https://horizonpulse.dev/#catalog",
    "text": "Catalog"
   }
  ],
  "title": "Horizon Pulse",
  "fields": {
   "links": [
    "https://horizonpulse.dev/",
    "https://horizonpulse.dev/#catalog",
    "https://horizonpulse.dev/#how"
   ],
   "heading": "Pay-per-call APIsbuilt for agents."
  },
  "finalUrl": "https://horizonpulse.dev/",
  "headings": [
   {
    "text": "Pay-per-call APIs built for agents.",
    "level": 1
   },
   {
    "text": "Thirteen routes. One protocol.",
    "level": 2
   }
  ],
  "language": "en",
  "elapsedMs": 120,
  "textSample": "# Horizon Pulse\n\nLive on Base · x402 v2\n# Pay-per-call APIs\n built for agents.\n\n13 routes for market data, the web and d…",
  "description": "Pay-per-call APIs for AI agents via x402 on Base: web fetch, HTTP proxy, page extract, and crypto market data.",
  "fieldErrors": {},
  "matchedFields": 2,
  "requestedFields": 2
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/horizonpulse-html-extractor-5aa58644/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from horizonpulse.dev](https://www.zero.xyz/host/horizonpulse.dev/llms.txt)
