# Anokha Web Tools – PDF Text Extractor

> Anokha Web Tools – PDF Text Extractor is a paid API for AI agents from tools.anokha.space, paid per call via x402, $0.006/call, status unknown (last checked 2026-10-02).

Fetches a publicly accessible PDF by URL and returns its extracted text content, page count, and title

## Facts

- Endpoint: GET https://tools.anokha.space/v1/pdf?utm_source=zero.xyz
- Price: $0.006/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/anokha-web-tools-pdf-text-extractor-cbee33fa
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_QdY5fh8hGv9t_OfmkEOKu

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability anokha-web-tools-pdf-text-extractor-cbee33fa
```

Example prompt: Can you grab the text from this PDF — https://example.com/report.pdf — and tell me what it says, including how many pages it has?

## When to prefer this

Use this endpoint when you need to read the textual content of a publicly hosted PDF without downloading it locally. It is ideal for AI agents that need to ingest, summarize, or analyze PDF documents on-the-fly, especially in a pay-per-call serverless context where no API key management is desired and payment is handled via x402/USDC on Base.

## Known failure modes

- URL is not publicly accessible or returns a non-200 status — extraction fails
- URL does not point to a valid PDF file — parse error returned
- PDF is scanned/image-based with no embedded text — empty or minimal text returned per page
- PDF is password-protected — extraction fails
- Very large PDFs may time out or return partial results

## How this service works

Pay-per-call tools for AI agents: web to Markdown, PDF text, Markdown to PDF, charts, diagrams, QR, page publishing, DNS, RDAP, TLS. x402, USDC on Base, no API key.

## Output

A JSON object containing the original URL, the total number of pages, the document title, and an array of page objects each with a page number and the extracted text from that page.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "required": [
  "input"
 ],
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "type",
    "method"
   ],
   "properties": {
    "type": {
     "type": "string",
     "const": "http"
    },
    "method": {
     "enum": [
      "GET",
      "HEAD",
      "DELETE"
     ],
     "type": "string"
    },
    "queryParams": {
     "type": "object",
     "required": [
      "url"
     ],
     "properties": {
      "url": {
       "type": "string",
       "description": "Public http(s) URL"
      }
     }
    }
   },
   "additionalProperties": false
  },
  "output": {
   "type": "object",
   "required": [
    "type"
   ],
   "properties": {
    "type": {
     "type": "string"
    },
    "example": {
     "type": "object"
    }
   }
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "url": "...",
  "text": [
   {
    "page": 1,
    "text": "..."
   }
  ],
  "pages": 12,
  "title": "..."
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/anokha-web-tools-pdf-text-extractor-cbee33fa/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from tools.anokha.space](https://www.zero.xyz/host/tools.anokha.space/llms.txt)
