# x402-markdown-extractor PDF to Markdown

> x402-markdown-extractor PDF to Markdown is a paid API for AI agents from x402.valkyry.fr, paid per call via x402, $0.003/call, status unknown (last checked 2026-10-01).

Converts a publicly accessible PDF URL into extracted markdown text, returning page count and word count

## Facts

- Endpoint: POST https://x402.valkyry.fr/pdf-to-markdown?utm_source=zero.xyz
- Price: $0.003/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-01
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/x402-markdown-extractor-pdf-to-markdown-3a977789
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_ROAEALT5mgkLxkrH5PeRw

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability x402-markdown-extractor-pdf-to-markdown-3a977789 -d '<json body>'
```

Example prompt: Can you extract the text from this PDF and give it to me as markdown? The file is at https://example.com/report.pdf

## When to prefer this

Use this endpoint when you need to extract the full text content of a remotely hosted PDF as markdown, especially for AI pipelines that need to process or summarize document content. It is ideal for single-document extraction where a public URL is available and you want structured markdown output with metadata like page and word count. Prefer this over general web scrapers when your target is specifically a PDF file.

## Known failure modes

- Private, loopback, or metadata IP addresses are rejected with an error — URL must be publicly reachable
- Non-PDF URLs or URLs that return non-PDF content will fail to parse
- Very large PDFs may time out or return partial content
- Scanned/image-only PDFs with no embedded text will return empty or minimal markdown
- Payment failure via x402 protocol returns HTTP 402 before processing occurs
- Malformed or non-absolute URLs are rejected by input validation

## How this service works

Fetch a PDF by URL and extract its text as Markdown, with page count and word count. Lets AI agents read PDF documents, reports and papers. (Text-based PDFs; scanned/image-only PDFs are not OCR'd.)

## Output

A JSON object containing the original URL, total page count (integer), extracted markdown text of the full PDF content, and a word count integer.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "url": {
   "type": "string",
   "format": "uri",
   "description": "Absolute http(s) URL. Must be publicly reachable; private/loopback/metadata addresses are rejected."
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "url": "https://example.com/doc.pdf",
  "pages": 12,
  "markdown": "Extracted text…",
  "wordCount": 3400
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/x402-markdown-extractor-pdf-to-markdown-3a977789/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from x402.valkyry.fr](https://www.zero.xyz/host/x402.valkyry.fr/llms.txt)
