# Crawlspur X-Robots-Tag Checker for News Documents

> Crawlspur X-Robots-Tag Checker for News Documents is a paid API for AI agents from crawlspur.halowerk.com, paid per call via x402, $0.001/call, status unknown (last checked 2026-09-14).

Checks the X-Robots-Tag HTTP header value on a news document URL and returns the crawl/indexing directive found.

## Facts

- Endpoint: POST https://crawlspur.halowerk.com/v1/pruef/crawlspur/x-robots/news-dokument
- Price: $0.001/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-14
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/crawlspur-x-robots-tag-checker-for-news-documents-f11f4a47
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_YHfjgQ7mNYFISwZ2KlZX5

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability crawlspur-x-robots-tag-checker-for-news-documents-f11f4a47 -d '<json body>'
```

Example prompt: Check the X-Robots-Tag HTTP header on this news article and tell me what crawl directive it returns: https://www.example-news.com/2024/01/breaking-story.html

## When to prefer this

Use this endpoint when you specifically need to inspect the X-Robots-Tag HTTP header on a news document — distinct from meta robots tags or robots.txt directives. Prefer this over generic HTTP header tools when the target is a news page and you need a structured, SEO-focused interpretation of the X-Robots-Tag value. Choose this over the sibling 'bild' (image) or 'robots-datei' endpoints when the object type is specifically a news article or news document URL.

## Known failure modes

- Target URL is not publicly accessible or returns a non-200 HTTP status
- URL does not point to a news document, returning unexpected content type
- X-Robots-Tag header is absent on the response (may return empty or null result)
- URL exceeds maximum length of 2048 characters
- Network timeout or DNS resolution failure for the target address
- Invalid input schema (missing required 'ziel' field)

## How this service works

Prüft am Objekt news-dokument den Befund Wert des X-Robots-Tags.

## Output

Returns the value of the X-Robots-Tag HTTP response header found on the specified news document, indicating any crawl or indexing directives (e.g. noindex, nofollow, none, unavailable_after) that search engine bots would receive when visiting the page.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "ziel": {
   "oneOf": [
    {
     "type": "string",
     "maxLength": 2048,
     "minLength": 1,
     "description": "Öffentliche HTTP(S)-Adresse; bei DNS-Befunden Domain oder öffentliche IP."
    },
    {
     "type": "object",
     "required": [
      "adresse"
     ],
     "properties": {
      "adresse": {
       "type": "string",
       "maxLength": 2048,
       "minLength": 1
      },
      "selektor": {
       "type": "string",
       "maxLength": 240,
       "minLength": 1,
       "description": "CSS-Selektor: genau ein Objekt, sonst erster Treffer des Objektfilters."
      },
      "vergleich": {
       "type": "string",
       "maxLength": 2048,
       "description": "Öffentliche Vergleichsadresse für Link-, Sitemap- oder Sprachbefunde."
      },
      "user_agent": {
       "type": "string",
       "pattern": "^[A-Za-z0-9_-]{1,80}$"
      },
      "dkim_selektor": {
       "type": "string",
       "pattern": "^[A-Za-z0-9_-]{1,63}$"
      }
     },
     "additionalProperties": false
    }
   ]
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/crawlspur-x-robots-tag-checker-for-news-documents-f11f4a47/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from crawlspur.halowerk.com](https://www.zero.xyz/host/crawlspur.halowerk.com/llms.txt)
