# Crawlspur robots.txt Crawl Permission Checker

> Crawlspur robots.txt Crawl Permission Checker is a paid API for AI agents from crawlspur.halowerk.com, paid per call via x402, $0.001/call, status unknown (last checked 2026-09-14).

Checks whether a given URL is allowed to be crawled according to the site's robots.txt file.

## Facts

- Endpoint: POST https://crawlspur.halowerk.com/v1/pruef/crawlspur/robots-erlaubt/robots-datei
- Price: $0.001/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-14
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/crawlspur-robots-txt-crawl-permission-checker-5d780d4a
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_8I-ump5PF9hOTM6ZuDIm6

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability crawlspur-robots-txt-crawl-permission-checker-5d780d4a -d '<json body>'
```

Example prompt: Can you check whether https://example.com/blog/my-article is allowed to be crawled according to its robots.txt file, using the user-agent Googlebot?

## When to prefer this

Use this endpoint when you need to programmatically verify whether a specific URL is permitted for crawling under robots.txt rules before initiating automated access, during SEO audits, or when building compliant web scrapers. Prefer this over manual inspection when checking many URLs or integrating crawl-permission checks into automated pipelines.

## Known failure modes

- Target URL is unreachable or returns a non-200 HTTP status
- robots.txt file is missing or returns an error, making permission indeterminate
- Invalid or malformed URL provided in the 'ziel' field
- User-agent pattern does not match any robots.txt rule, defaulting to allow
- Network timeout when fetching the robots.txt file

## How this service works

Prüft am Objekt robots-datei den Befund Freigabe durch robots.txt.

## Output

Returns a finding indicating whether the target URL is permitted (allowed or disallowed) by the site's robots.txt file, based on the applicable crawl directives for the specified or default user-agent.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "ziel": {
   "oneOf": [
    {
     "type": "string",
     "maxLength": 2048,
     "minLength": 1,
     "description": "Öffentliche HTTP(S)-Adresse; bei DNS-Befunden Domain oder öffentliche IP."
    },
    {
     "type": "object",
     "required": [
      "adresse"
     ],
     "properties": {
      "adresse": {
       "type": "string",
       "maxLength": 2048,
       "minLength": 1
      },
      "selektor": {
       "type": "string",
       "maxLength": 240,
       "minLength": 1,
       "description": "CSS-Selektor: genau ein Objekt, sonst erster Treffer des Objektfilters."
      },
      "vergleich": {
       "type": "string",
       "maxLength": 2048,
       "description": "Öffentliche Vergleichsadresse für Link-, Sitemap- oder Sprachbefunde."
      },
      "user_agent": {
       "type": "string",
       "pattern": "^[A-Za-z0-9_-]{1,80}$"
      },
      "dkim_selektor": {
       "type": "string",
       "pattern": "^[A-Za-z0-9_-]{1,63}$"
      }
     },
     "additionalProperties": false
    }
   ]
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/crawlspur-robots-txt-crawl-permission-checker-5d780d4a/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from crawlspur.halowerk.com](https://www.zero.xyz/host/crawlspur.halowerk.com/llms.txt)
