# PDF Document Extractor

> PDF Document Extractor is a paid API for AI agents from dt0ur.online, paid per call via x402, $0.05/call, status unknown (last checked 2026-09-30).

Parses multi-page PDF documents from a public URL and returns structured text content along with page counts.

## Facts

- Endpoint: POST https://dt0ur.online/api/documents/pdf-extract?utm_source=zero.xyz
- Price: $0.05/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-30
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/pdf-document-extractor-a310025d
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_fV3BK92Shp7bYEL0utqW1

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability pdf-document-extractor-a310025d -d '<json body>'
```

Example prompt: Can you extract all the text from this PDF for me? Here's the URL: https://example.com/reports/annual-report-2024.pdf — I need the full text content and the page count.

## When to prefer this

Choose this endpoint when you have a publicly accessible URL pointing to a PDF document and need to extract its textual content programmatically. It is ideal for document workflows, contract review pipelines, research aggregation, or any scenario where a PDF needs to be converted to machine-readable text without OCR. Prefer it over general scraping tools when the content is specifically PDF-formatted and multi-page structure with page counts matters.

## Known failure modes

- PDF URL is not publicly accessible or returns a non-200 HTTP status
- URL does not point to a valid PDF file (e.g. HTML page or image)
- PDF is password-protected or encrypted and cannot be parsed
- PDF contains only scanned images with no embedded text (requires OCR instead)
- Network timeout fetching the remote PDF document
- Malformed or corrupted PDF file that cannot be parsed

## How this service works

Parses multi-page PDF documents from URLs, returning structured text content and page counts.

## Output

Returns the structured text content extracted from each page of the PDF document along with the total page count, allowing agents to read, search, summarize, or further process the document's contents.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "pdf_url": {
   "type": "string",
   "format": "uri",
   "description": "Public URL to the PDF document"
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/pdf-document-extractor-a310025d/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from dt0ur.online](https://www.zero.xyz/host/dt0ur.online/llms.txt)
