# x402-markdown-extractor

> x402-markdown-extractor is a paid API for AI agents from x402.valkyry.fr, paid per call via x402, $0.002/call, status unknown (last checked 2026-09-30).

Fetches a publicly reachable URL and returns its main content as clean Markdown, including title, word count, and fetch timestamp.

## Facts

- Endpoint: POST https://x402.valkyry.fr/extract?utm_source=zero.xyz
- Price: $0.002/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-30
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/x402-markdown-extractor-6324e6c7
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_I3_Lm2jVW_ioYIxBENFYh

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability x402-markdown-extractor-6324e6c7 -d '<json body>'
```

Example prompt: Can you grab the content of https://example.com/some-article and give me the full text as clean markdown so I can read and summarize it?

## When to prefer this

Use this endpoint when an AI agent needs to retrieve and read the human-readable content of a specific public URL — particularly for summarization, Q&A, or document ingestion workflows — and when you want a pay-per-call model (USDC on Base via x402) rather than a subscription scraping service.

## Known failure modes

- Private, loopback, or metadata addresses rejected with an error
- URL unreachable or returns non-2xx HTTP status
- Page has no extractable body content, returning empty or minimal markdown
- Malformed or non-URI input rejected with validation error
- Payment failure via x402 protocol prevents the call from executing

## How this service works

Turn any web page into clean, LLM-ready Markdown. Fetches the URL, isolates the main article with Mozilla Readability (stripping nav, ads and boilerplate), and returns tidy Markdown with the page title and word count — built for AI agents, RAG pipelines, web scraping and content ingestion.

## Output

A JSON object containing the original URL, the page title, the full body text rendered as clean Markdown, the UTC timestamp when the page was fetched, and the word count of the extracted content.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "url": {
   "type": "string",
   "format": "uri",
   "description": "Absolute http(s) URL. Must be publicly reachable; private/loopback/metadata addresses are rejected."
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "url": "https://example.com/some-article",
  "title": "Some Article",
  "markdown": "# Some Article\n\nClean body text...",
  "fetchedAt": "2026-05-28T12:34:56.789Z",
  "wordCount": 842
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/x402-markdown-extractor-6324e6c7/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from x402.valkyry.fr](https://www.zero.xyz/host/x402.valkyry.fr/llms.txt)
