# Vextorium Web Markdown Extractor

> Vextorium Web Markdown Extractor is a paid API for AI agents from api.vextorium.com, paid per call via x402, $0.0045/call, status unknown (last checked 2026-10-02).

Fetches any public web page or PDF by URL and returns its main content as clean, LLM-ready markdown plus metadata (title, description, author, date, language, canonical URL), with boilerplate removed.

## Facts

- Endpoint: POST https://api.vextorium.com/web-markdown-pago?utm_source=zero.xyz
- Price: $0.0045/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/vextorium-web-markdown-extractor-cc4038d0
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_mhdJ81JtsilA5YkliMSS-

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability vextorium-web-markdown-extractor-cc4038d0 -d '<json body>'
```

Example prompt: Can you fetch the article at https://example.com/news/article-123 and give me its main content as clean markdown — strip out the menus, ads, and footers — and also return the title, author, and publication date?

## When to prefer this

Choose this endpoint when you need clean, boilerplate-free markdown from any public web page or PDF, especially when feeding content into an LLM context window. It is preferable over raw HTTP fetching because it removes menus, footers, and cookie banners automatically, and over general-purpose scrapers because it outputs structured LLM-ready markdown with rich metadata in a single call. Best for article reading, document ingestion, and research pipelines where content quality and format matter.

## Known failure modes

- URL is not publicly accessible or requires authentication — returns error or empty content
- SSRF-blocked URL (private IPs, localhost) — rejected by safety filter
- PDF parsing fails for scanned/image-only PDFs with no text layer
- Content too long and truncated at max_caracteres limit — recortado flag set to true
- Timeout on slow or unresponsive servers
- Paywalled content returns only teaser text
- Invalid URL format — returns validation error

## How this service works

Read any public web page or PDF and get its main content as clean, LLM-ready markdown (headings, lists, links, tables, code) plus metadata: title, description, author, date, language, canonical URL. Menus, cookie banners and footers removed. Optional list of links. SSRF-safe fetching.

## Output

Returns a JSON object with: markdown-formatted main content (headings, lists, tables, code, links), word count, character count, whether the content was truncated, the final resolved URL, content type, metadata object (title, description, canonical URL, language, site name, featured image), optional array of links found on the page, and any warnings encountered during fetch.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "required": [
  "input"
 ],
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "type",
    "method",
    "bodyType",
    "body"
   ],
   "properties": {
    "body": {
     "required": [
      "url"
     ],
     "properties": {
      "url": {
       "type": "string",
       "description": "Full http(s) URL of a web page or PDF"
      },
      "max_caracteres": {
       "type": "integer",
       "maximum": 100000,
       "minimum": 500,
       "description": "Maximum markdown length (500-100000, default 20000)"
      },
      "incluir_enlaces": {
       "type": "boolean",
       "description": "Also return the list of links found (default false)"
      }
     }
    },
    "type": {
     "type": "string",
     "const": "http"
    },
    "method": {
     "enum": [
      "POST"
     ],
     "type": "string"
    },
    "bodyType": {
     "enum": [
      "json",
      "form-data",
      "text"
     ],
     "type": "string"
    }
   },
   "additionalProperties": false
  },
  "output": {
   "type": "object",
   "required": [
    "type"
   ],
   "properties": {
    "type": {
     "type": "string"
    },
    "example": {
     "type": [
      "object",
      "null"
     ],
     "properties": {
      "url": {
       "type": [
        "string",
        "null"
       ]
      },
      "tipo": {
       "type": [
        "string",
        "null"
       ]
      },
      "avisos": {
       "type": [
        "array",
        "null"
       ],
       "items": {
        "type": [
         "null",
         "string"
        ]
       }
      },
      "markdown": {
       "type": [
        "string",
        "null"
       ]
      },
      "palabras": {
       "type": [
        "number",
        "null"
       ]
      },
      "metadatos": {
       "type": [
        "object",
        "null"
       ],
       "properties": {
        "sitio": {
         "type": [
          "string",
          "null"
         ]
        },
        "idioma": {
         "type": [
          "string",
          "null"
         ]
        },
        "imagen": {
         "type": [
          "string",
          "null"
         ]
        },
        "titulo": {
         "type": [
          "string",
          "null"
         ]
        },
        "canonica": {
         "type": [
          "string",
          "null"
         ]
        },
        "descripcion": {
         "type": [
          "string",
          "null"
 
… (truncated)
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/vextorium-web-markdown-extractor-cc4038d0/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from api.vextorium.com](https://www.zero.xyz/host/api.vextorium.com/llms.txt)
