# Aayat AI Web Crawler — URL to Markdown

> Aayat AI Web Crawler — URL to Markdown is a paid API for AI agents from aayatai.com, paid per call via x402, $0.03/call, status unknown (last checked 2026-10-02).

Crawls a website or section of it starting from a given URL and returns up to 10 pages as clean Markdown, respecting robots.txt and staying within an optional path prefix.

## Facts

- Endpoint: GET https://aayatai.com/crawl?utm_source=zero.xyz
- Price: $0.03/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/aayat-ai-web-crawler-url-to-markdown-6d282a33
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_WRbCOryi17-JRSU4IJaFh

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability aayat-ai-web-crawler-url-to-markdown-6d282a33
```

Example prompt: Can you crawl https://stripe.com/docs/api/ and give me the content of up to 5 pages under /docs/api/ as clean Markdown?

## When to prefer this

Choose this endpoint when you need to ingest multi-page website content as clean Markdown in a single API call, especially for documentation sites, blogs, or company knowledge bases. It is ideal for RAG pipelines, LLM context loading, or competitive research where you want structured text from several related pages under a common URL path. Prefer it over single-page scrapers when you need up to 10 pages at once and want automatic sitemap discovery, robots.txt compliance, and link-following built in.

## Known failure modes

- Invalid or unreachable start URL returns an error or empty pages array
- robots.txt disallows crawling, resulting in all pages skipped
- Site has no sitemap and no followable links, returning only the start page
- maxPages set to 0 or above 10 triggers validation error
- Path prefix filters out all discovered pages, returning empty results
- Non-HTML content (PDFs, images) is skipped and listed in the skipped array
- Network timeouts on slow or firewalled sites may result in partial or empty results

## How this service works

Turn a small website, or one section of it, into clean Markdown: starts from your URL, finds pages from the sitemap or by following links on the same site, obeys robots.txt, and returns up to 10 pages (title, Markdown, address) in one call. Great for docs, blogs and company sites. ?url=https://example.com/docs/&maxPages=10&path=/docs/

## Output

A JSON object containing: the start URL, number of pages returned, how pages were discovered (sitemap or link-following), robots.txt status, a list of page objects each with URL, title, Markdown content, character count, and truncation flag, a list of skipped URLs with reasons, a completeness flag indicating if more pages existed than were returned, and a crawl timestamp.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "required": [
  "input"
 ],
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "type",
    "method"
   ],
   "properties": {
    "type": {
     "type": "string",
     "const": "http"
    },
    "method": {
     "enum": [
      "GET"
     ],
     "type": "string"
    },
    "queryParams": {
     "type": "object",
     "required": [
      "url"
     ],
     "properties": {
      "url": {
       "type": "string",
       "format": "uri",
       "maxLength": 2048,
       "description": "Start page, e.g. https://example.com/docs/."
      },
      "path": {
       "type": "string",
       "maxLength": 300,
       "description": "Only crawl pages under this path (default: the start page's folder)."
      },
      "maxPages": {
       "type": "integer",
       "default": 5,
       "maximum": 10,
       "minimum": 1,
       "description": "Most pages to return (1-10)."
      }
     }
    }
   },
   "additionalProperties": false
  },
  "output": {
   "type": "object",
   "required": [
    "type"
   ],
   "properties": {
    "type": {
     "type": "string"
    },
    "example": {
     "type": "object",
     "required": [
      "start",
      "pages",
      "count",
      "discoveredFrom",
      "robots",
      "complete"
     ],
     "properties": {
      "count": {
       "type": "integer"
      },
      "pages": {
       "type": "array",
       "items": {
        "type": "object",
        "properties": {
         "url": {
          "type": "string"
         },
         "chars": {
          "type": "integer"
         },
         "title": {
          "type": "string"
         },
         "markdown": {
          "type": "string"
         },
         "truncated": {
          "type": "boolean"
         }
        }
       }
      },
      "start": {
       "type": "string"
      },
      "trust": {
       "type": "object",
       "description": "Third-party text, cleaned: read trust.notice; removed = what we stripped."
      },
      "robots": {
       "type": "string",
       "description": "robots.txt status: ok, none or unreachable."
      },
      "skipped": {
       "type": "array",
       "items": {
        "type": "object"
       },
       "description": "Pages not read and why (robots.txt, error, not HTML)."
      },
      "complete": {
       "type": "boolean",
       "description": "False if more matching pages were found than returned."
      },
      "crawled
… (truncated)
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "count": 1,
  "pages": [
   {
    "url": "https://example.com/",
    "chars": 64,
    "title": "Example Domain",
    "markdown": "# Example Domain\n\nThis domain is for use in illustrative examples.",
    "truncated": false
   }
  ],
  "start": "https://example.com/",
  "robots": "ok",
  "skipped": [
   {
    "url": "https://example.com/admin",
    "reason": "robots.txt"
   }
  ],
  "complete": true,
  "crawledAt": "2026-09-28T12:00:00.000Z",
  "discoveredFrom": "links"
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/aayat-ai-web-crawler-url-to-markdown-6d282a33/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from aayatai.com](https://www.zero.xyz/host/aayatai.com/llms.txt)
