# AgentTools PDF to Text Extractor

> AgentTools PDF to Text Extractor is a paid API for AI agents from agenttools-hub.vercel.app, paid per call via x402, $0.003/call, status unknown (last checked 2026-10-02).

Fetches a PDF by URL or accepts it as base64 and extracts all readable text with page count using deterministic parsing (no LLM).

## Facts

- Endpoint: GET https://agenttools-hub.vercel.app/api/v1/dev/pdf-to-text?utm_source=zero.xyz
- Price: $0.003/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/agenttools-pdf-to-text-extractor-e0ea3a48
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_YJ3Re96bxt4YDpQ4PCNqa

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability agenttools-pdf-to-text-extractor-e0ea3a48
```

Example prompt: Can you extract all the text from this PDF for me? Here's the link: https://example.com/report.pdf — just pull out all the readable text.

## When to prefer this

Use this endpoint when an agent needs to read the contents of a text-based PDF — either by URL or as base64 — and requires deterministic, faithful extraction with no hallucination risk. Prefer this over LLM-based PDF tools when accuracy and verbatim reproduction are critical. Not suitable for scanned/image PDFs that require OCR.

## Known failure modes

- URL is unreachable or returns non-PDF content
- PDF is a scanned image (no embedded text layer) — returns empty or minimal text
- SSRF guard blocks internal/private IP URLs
- base64 input is malformed or not a valid PDF
- maxChars limit truncates output unexpectedly
- PDF is password-protected or encrypted

## How this service works

Fetch a PDF (by URL, SSRF-guarded) or accept it as base64 and extract all readable text with page count. Deterministic parsing — no language model, so nothing is invented. Text-based PDFs only; scanned images need OCR (not included). Use this when an agent needs to extract readable text from any PDF — agents can't read PDFs.

## Output

Returns all extracted plain text from the PDF along with the total page count. Output is deterministic — no language model is involved, so only text actually present in the PDF is returned. Works on text-based PDFs only; scanned image PDFs require OCR which this endpoint does not provide.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "required": [
  "input"
 ],
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "type",
    "method"
   ],
   "properties": {
    "type": {
     "type": "string",
     "const": "http"
    },
    "method": {
     "enum": [
      "GET"
     ],
     "type": "string"
    },
    "queryParams": {
     "type": "object",
     "required": [],
     "properties": {
      "url": {
       "type": "string",
       "description": "Or provide base64 instead"
      },
      "base64": {
       "type": "string",
       "description": "Base64 content"
      },
      "maxChars": {
       "type": "number",
       "description": "Max text length"
      }
     }
    }
   },
   "additionalProperties": false
  },
  "output": {
   "type": "object",
   "required": [
    "type"
   ],
   "properties": {
    "type": {
     "type": "string"
    },
    "example": {
     "type": "object"
    }
   }
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/agenttools-pdf-to-text-extractor-e0ea3a48/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from agenttools-hub.vercel.app](https://www.zero.xyz/host/agenttools-hub.vercel.app/llms.txt)
