# x402.shizu.me PDF Text Extractor

> x402.shizu.me PDF Text Extractor is a paid API for AI agents from x402.shizu.me, paid per call via x402, $0.005/call, status unknown (last checked 2026-09-29).

Fetches a public PDF by URL and returns its extracted text content as structured JSON with page count and word count

## Facts

- Endpoint: GET https://x402.shizu.me/pdf?utm_source=zero.xyz
- Price: $0.005/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-29
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/x402-shizu-me-pdf-text-extractor-e8e71b67
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_fjWai3nwBCnM52_XYJA9h

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability x402-shizu-me-pdf-text-extractor-e8e71b67
```

Example prompt: Can you pull the text out of this PDF for me — https://example.com/report.pdf — I want to read its contents?

## When to prefer this

Use this endpoint when you need to extract raw text from a publicly accessible PDF file given only its URL, especially when feeding PDF content into an LLM pipeline. Ideal for single-document extraction under 3 MB and 50 pages where you need structured JSON output including page and word counts.

## Known failure modes

- URL is not a valid public HTTP/HTTPS address — returns error
- PDF exceeds 3 MB file size limit — returns error or truncated result
- PDF has more than 50 pages — only first 50 pages are processed
- URL does not point to a PDF file — parsing fails
- PDF is password-protected or encrypted — text extraction fails
- Network timeout fetching the remote PDF — returns error

## How this service works

Fetch a PDF by URL and extract its text as JSON. Input: url (public http/https pointing to a PDF). Returns {url, pages, word_count, text}. Bounded to the first 50 pages; 3 MB max.

## Output

Returns a JSON object with the original URL, number of pages processed (up to 50), total word count, and the full extracted text from the PDF

## Request schema (JSON Schema)

```json
{
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "method"
   ],
   "properties": {
    "method": {
     "const": "GET"
    },
    "queryParams": {
     "type": "object",
     "required": [
      "url"
     ],
     "properties": {
      "url": {
       "type": "string"
      }
     }
    }
   }
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "url": "https://example.com/file.pdf",
  "text": "Extracted text content of the PDF...",
  "pages": 3,
  "word_count": 412
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/x402-shizu-me-pdf-text-extractor-e8e71b67/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from x402.shizu.me](https://www.zero.xyz/host/x402.shizu.me/llms.txt)
