# AI Data Tools - Scraped Data Cleaner

> AI Data Tools - Scraped Data Cleaner is a paid API for AI agents from www.aidatatools.dev, paid per call via x402, $0.04/call, status unknown (last checked 2026-09-13).

Deterministically repairs and normalizes scraped data (JSON objects, arrays, CSV, or raw text) and returns the same data shape, cleaned.

## Facts

- Endpoint: POST https://www.aidatatools.dev/api/clean
- Price: $0.04/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-13
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/ai-data-tools-scraped-data-cleaner-111f2324
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_kxB4AHeKwEpz89XPWUCxc

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability ai-data-tools-scraped-data-cleaner-111f2324 -d '<json body>'
```

Example prompt: I just scraped a batch of product records and some have garbled prices like '12,99', weird special characters in titles, and inconsistent SKU formatting — can you run them through the data cleaner and give me back the fixed records in the same JSON shape?

## When to prefer this

Choose this endpoint when you have scraped data with deterministic, rule-based repair needs — malformed prices, encoding issues, special character corruption, or inconsistent formatting — and you need the exact same data shape returned. Prefer it over ML-based enrichment tools when reproducibility and determinism are critical, and over validation-only endpoints when you need the data actually fixed, not just flagged. It accepts JSON objects, arrays, CSV, and raw text, making it versatile for scraping pipelines.

## Known failure modes

- Unrecognizable input format that cannot be parsed — likely returns an error or unparseable response
- Input body missing required 'type', 'method', or 'bodyType' fields — returns 400-level validation error
- Extremely large payloads may time out or be rejected
- Ambiguous data where the 'correct' value is unclear may be repaired inconsistently
- Non-UTF-8 or binary content not representable as JSON/text may fail to parse

## How this service works

Deterministic post-scrape data cleaner and quality gate for AI agents. Three tiers over one engine, no LLM anywhere: the same input always produces byte-identical output.

**CLEAN** (`POST /api/clean`, $0.04) - post the raw output of a scrape, get the REPAIRED data back as the response body: residual HTML stripped, mojibake decoded ("CafÃ©" -> "Café"), invisible characters removed, non-breaking spaces normalised, values trimmed, across nested objects and arrays. It repairs how data was ENCODED and never what it SAYS: a negative price or a failed extraction ("captcha", "access denied") is reported, never rewritten or deleted. Call it after every extraction run - a verdict is cached per source, but dirt is produced fresh by every run.

**CLEAN + AUDIT** (`POST /api/clean/audit`, $0.12) - identical repaired data plus a complete, replayable, reversible ledger of every transformation, with a replay_id and input/output SHA-256. Applying the ledger in reverse reconstructs the input byte for byte.

**VERDICT** (`POST /api`, $0.01) - score + exact facts + a RELIABLE / USABLE_WITH_CLEANING / UNRELIABLE judgement, for deciding whether to trust a source at all. Facts-only signals (price_divergence, text_cleanliness, a robust MAD cross-check) report alongside without moving the score.

What is repaired automatically, what needs an explicit opt-in, and what is only ever reported is published in full at `GET /api/clean` - machine-readable, and auditable before you pay. Paid via x402: no account, no API key, no signup.

## Output

Returns the same data shape as the input (JSON array of objects, single object, CSV string, or raw text) with fields repaired: normalized prices, corrected special characters, standardized encodings, and fixed formatting inconsistencies — a drop-in replacement for the original scraped data.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "required": [
  "input"
 ],
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "type",
    "method",
    "bodyType",
    "body"
   ],
   "properties": {
    "body": {
     "type": "object",
     "description": "The scraped data to process. A JSON array of objects, a single object, a CSV string or raw text are all accepted -- the endpoint detects the shape and gives the same shape back."
    },
    "type": {
     "type": "string",
     "const": "http"
    },
    "method": {
     "enum": [
      "POST",
      "PUT",
      "PATCH"
     ],
     "type": "string"
    },
    "bodyType": {
     "enum": [
      "json",
      "form-data",
      "text"
     ],
     "type": "string"
    }
   },
   "additionalProperties": false
  },
  "output": {
   "type": "object",
   "required": [
    "type"
   ],
   "properties": {
    "type": {
     "type": "string"
    },
    "example": {
     "type": "object"
    }
   }
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "sku": "B0C1",
  "price": "12,99",
  "title": "Café Table Lamp & Shade"
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/ai-data-tools-scraped-data-cleaner-111f2324/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from www.aidatatools.dev](https://www.zero.xyz/host/www.aidatatools.dev/llms.txt)
