# Web Article & PDF to Markdown Converter

> Web Article & PDF to Markdown Converter is a paid API for AI agents from x402.donnyautomation.com, paid per call via x402, $0.01/call, status unknown (last checked 2026-09-28).

Fetches any public URL (article or PDF) and returns clean Markdown with title, byline, site name, excerpt, and word count using Firefox Readability extraction.

## Facts

- Endpoint: GET https://x402.donnyautomation.com/markdown?utm_source=zero.xyz
- Price: $0.01/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-28
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/web-article-pdf-to-markdown-converter-fe561683
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_Iztjr3-W7bLhX4JUgYkqu

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability web-article-pdf-to-markdown-converter-fe561683
```

Example prompt: Can you fetch this article at https://en.wikipedia.org/wiki/Artificial_intelligence and give me the clean markdown text along with the title and word count?

## When to prefer this

Choose this endpoint when you need reliable, honest article or PDF text extraction from a public URL with clean Markdown output and rich metadata (title, byline, excerpt, word count). It is ideal for AI research pipelines, summarizers, and content ingestion workflows that need to trust the output — unlike generic scrapers that silently return empty or partial pages, this endpoint explicitly signals failures like image-only PDFs or non-extractable app shells. Prefer it over raw HTML fetchers when your downstream task is text-based (LLM summarization, RAG ingestion, content analysis).

## Known failure modes

- Image-only PDFs return no_text_layer status with no markdown body
- Client-rendered single-page apps return not_extractable when no static HTML is available
- Private or internal hostnames are refused with an error
- URLs exceeding the 8 MB fetch cap or 400K character output cap return an error
- Paywalled or login-protected pages may return incomplete or no content
- Invalid or malformed URLs return an error

## How this service works

Fetch a public article or PDF and return clean Markdown plus title, byline, siteName, excerpt and wordCount. HTML is extracted with Firefox reader-mode rules; PDFs return their text layer. Requires ?url=<public http(s) URL>. Errors: 400 missing_url|bad_url, 403 blocked_private (private and internal hosts refused, every redirect hop re-checked), 413 too_large above 8 MB, 422 no_text_layer for scanned PDFs or not_extractable for app shells, 504 fetch_timeout. To FIND urls, use /search.

## Output

Returns a structured response containing: the full article or PDF content as clean Markdown text, the page title, byline/author, site name, a short excerpt, and word count (or page count for PDFs). For image-only PDFs, returns a no_text_layer status. For JavaScript-rendered app shells that cannot be extracted, returns not_extractable — never silently returns an empty or misleading result.

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "ts": "2026-08-01T00:00:00.000Z",
  "url": "https://en.wikipedia.org/wiki/Markdown",
  "title": "Markdown",
  "byline": null,
  "format": "article",
  "excerpt": "Markdown is a lightweight markup language…",
  "finalUrl": "https://en.wikipedia.org/wiki/Markdown",
  "markdown": "# Markdown\n\nMarkdown is a lightweight markup language…",
  "siteName": "Wikipedia",
  "truncated": false,
  "wordCount": 3204
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/web-article-pdf-to-markdown-converter-fe561683/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from x402.donnyautomation.com](https://www.zero.xyz/host/x402.donnyautomation.com/llms.txt)
