# robots.txt Parser: Rules, Sitemaps, AI-Crawler Blocks, and Path-Allowed Check

> robots.txt Parser: Rules, Sitemaps, AI-Crawler Blocks, and Path-Allowed Check is a paid API for AI agents from twin.unykorn.org, paid per call via x402, $0.001/call, status unknown (last checked 2026-10-01).

Fetches and parses a site's robots.txt, returning crawl rules, sitemaps, AI-crawler blocks, and whether a specific path is allowed for a given user-agent.

## Facts

- Endpoint: POST https://twin.unykorn.org/web/robots?utm_source=zero.xyz
- Price: $0.001/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-01
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/robots-txt-parser-rules-sitemaps-ai-crawler-blocks-and-path-allowed-a48f69e4
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_PJvlyrwnywzF_4WboUYpM

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability robots-txt-parser-rules-sitemaps-ai-crawler-blocks-and-path-allowed-a48f69e4 -d '<json body>'
```

Example prompt: Check robots.txt for example.com and tell me whether /blog/ is allowed for the default user-agent, and list any sitemaps and AI-crawler blocks you find.

## When to prefer this

Use this endpoint when an agent needs to programmatically determine crawl permissions for any website before scraping, to audit AI-crawler policies across sites, or to discover sitemap URLs. Prefer this over manual fetching because it handles parsing, normalization, and AI-specific block detection automatically.

## Known failure modes

- robots.txt not found (404) — site may not have one, returns empty or error
- network timeout if the target domain is unreachable
- malformed robots.txt on the target site may produce incomplete parsing
- rate limiting by the target domain when fetching robots.txt
- invalid domain or URL input causes a parameter error

## How this service works

robots.txt parser: rules, sitemaps, AI-crawler blocks, and is-this-path-allowed check — Genesis402 / UnyKorn Operator Network

## Output

Returns the parsed robots.txt contents including crawl rules per user-agent, disallow and allow directives, sitemap URLs, whether specific AI crawlers are blocked, and a boolean indicating if the requested path is permitted for the specified user-agent.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "params": {
   "type": "object",
   "properties": {
    "path": {
     "type": "string",
     "description": "optional path to check, e.g. /blog/"
    },
    "site": {
     "type": "string",
     "description": "required: domain or URL"
    },
    "user_agent": {
     "type": "string",
     "description": "optional, default *"
    }
   }
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "ok": true,
  "type": "robots-txt",
  "receipt": {
   "tx_hash": "0x<64hex>",
   "amount_usd": 0.001,
   "receipt_id": "g402-<16hex>"
  },
  "sources": [
   {
    "ok": true,
    "name": "<source>"
   }
  ],
  "limitations": "<text>",
  "generated_at": "<iso time>",
  "evidence_hash": "sha256:<64hex>"
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/robots-txt-parser-rules-sitemaps-ai-crawler-blocks-and-path-allowed-a48f69e4/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from twin.unykorn.org](https://www.zero.xyz/host/twin.unykorn.org/llms.txt)
