# Horizon Pulse PDF to Text Extractor

> Horizon Pulse PDF to Text Extractor is a paid API for AI agents from horizonpulse.dev, paid per call via x402, $0.02/call, status unknown (last checked 2026-10-01).

Fetches a public PDF URL and returns clean extracted text per page plus title and author metadata, using pdf.js text layer (no OCR), limited to 10MB, 50 pages, and 100K characters.

## Facts

- Endpoint: GET https://horizonpulse.dev/api/pdf?utm_source=zero.xyz
- Price: $0.02/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-01
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/horizon-pulse-pdf-to-text-extractor-1d6affbd
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_WK3dVJoTB0YDNmqAeV0_W

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability horizon-pulse-pdf-to-text-extractor-1d6affbd
```

Example prompt: Can you grab the text from this PDF — https://example.gov/report.pdf — and give me the content of the first 10 pages along with whatever title and author info is in the document?

## When to prefer this

Choose this endpoint when you have a public HTTP/HTTPS URL pointing to a real PDF with an embedded text layer (not a scanned image PDF) and need clean per-page text plus title/author metadata. It is ideal for processing research papers, reports, filings, and other text-layer PDFs up to 10MB and 50 pages. Prefer this over generic web scrapers when the source is specifically a PDF. Do not use for scanned/image PDFs (no OCR), password-protected files, or PDFs larger than 10MB.

## Known failure modes

- PDF exceeds 10MB size limit — request rejected or truncated
- PDF requires OCR (image-only scans) — no text extracted, unbilled
- PDF is password-protected or encrypted — cannot parse, unbilled
- URL does not point to a valid PDF — non-PDF response, unbilled
- SSRF-blocked URL (private IP ranges, internal hostnames) — request rejected
- PDF has more than 50 pages — only first 50 pages returned
- Text exceeds 100K character limit — output truncated at that boundary
- URL is unreachable or returns non-200 status — fetch error

## How this service works

Pay-per-call APIs for AI agents via x402 on Base: web fetch, HTTP proxy, page extract, and crypto market data.

## Output

Returns clean text extracted per page from the PDF, along with document-level metadata such as title and author. Output is structured by page, covering up to 50 pages and 100K characters. Files that are non-PDF, encrypted, or image-only (requiring OCR) are returned unbilled with an appropriate indication.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "required": [
  "input"
 ],
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "type",
    "method"
   ],
   "properties": {
    "type": {
     "type": "string",
     "const": "http"
    },
    "method": {
     "enum": [
      "GET",
      "HEAD",
      "DELETE"
     ],
     "type": "string"
    },
    "queryParams": {
     "type": "object",
     "required": [
      "url"
     ],
     "properties": {
      "url": {
       "type": "string",
       "description": "Absolute http(s) URL of a public PDF (required, max 10MB)."
      },
      "pages": {
       "type": "string",
       "description": "Max pages to return, integer 1-50 (default 50)."
      }
     }
    }
   },
   "additionalProperties": false
  },
  "output": {
   "type": "object",
   "required": [
    "type"
   ],
   "properties": {
    "type": {
     "type": "string"
    },
    "example": {
     "type": "object",
     "description": "Horizon Pulse PDF text extraction result. The example is a trimmed real response recorded from GET /api/demo/pdf at 2026-10-01T14:36Z UTC; live values differ on every call."
    }
   }
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "ok": true,
  "meta": {
   "title": "Horizon Pulse sample PDF",
   "author": "Horizon Pulse"
  },
  "bytes": 858,
  "pages": [
   {
    "page": 1,
    "text": "Horizon Pulse sample PDF\nPay-per-call APIs for AI agents over x402 (USDC on Base).\nThis file is the fixed input for the free /api/demo/pdf sample."
   }
  ],
  "finalUrl": "https://horizonpulse.dev/sample.pdf",
  "truncated": false,
  "totalPages": 1,
  "requestedUrl": "https://horizonpulse.dev/sample.pdf",
  "pagesReturned": 1
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/horizon-pulse-pdf-to-text-extractor-1d6affbd/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from horizonpulse.dev](https://www.zero.xyz/host/horizonpulse.dev/llms.txt)
