# Delx Commerce PDF OCR

> Delx Commerce PDF OCR is a paid API for AI agents from commerce.delx.ai, paid per call via x402, $0.01/call, status unknown (last checked 2026-10-02).

Extracts text from scanned or machine-printed PDFs using Tesseract 5 + Poppler, returning recognized text with confidence score and cryptographic hashes for verifiable delivery

## Facts

- Endpoint: POST https://commerce.delx.ai/api/v1/x402/pdf-ocr?utm_source=zero.xyz
- Price: $0.01/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/delx-commerce-pdf-ocr-9bdf08a7
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_TIXIHIjndBGqpYqLTQQLM

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability delx-commerce-pdf-ocr-9bdf08a7 -d '<json body>'
```

Example prompt: Can you extract all the text from this scanned PDF I'm sending you as a base64 string? Use page segmentation mode 6 and limit it to 10 pages — it's a machine-printed contract.

## When to prefer this

Choose this endpoint when you need fast, pay-per-call PDF OCR with no signup, verifiable cryptographic delivery receipts, and privacy guarantees (no document stored or sent to third parties). Ideal for AI agents operating autonomously with USDC micropayments via x402 protocol on Base or Solana. Prefer over subscription OCR services when you need per-result billing and auditability via SHA-256 hashes.

## Known failure modes

- PDF exceeds 4 MiB decoded size — request rejected
- PDF base64 string malformed or not valid PDF — parse error
- More than 10 pages requested — capped at maximum
- Character limit exceeded — text truncated with truncated:true flag
- Non-machine-printed or handwritten text — low confidence score or empty output

## How this service works

Pay-per-result APIs for agents. No signup. Exact price. Verifiable delivery. USDC on Base + Solana via x402.

## Output

A JSON object containing: the full recognized text (paginated with page markers), OCR confidence score (0–100), document ID, SHA-256 hashes of input and output for verifiability, page and word counts, language detected, model used (tesseract-5+poppler), truncation flag, and sale price in USDC. No document is stored or sent to third parties.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "psm": {
   "enum": [
    3,
    6,
    11
   ],
   "type": "integer",
   "default": 6,
   "description": "Tesseract page segmentation mode for machine-printed text."
  },
  "max_chars": {
   "type": "integer",
   "default": 200000,
   "maximum": 200000,
   "minimum": 1,
   "description": "Maximum UTF-8 characters returned; OCR remains bounded."
  },
  "max_pages": {
   "type": "integer",
   "default": 10,
   "maximum": 10,
   "minimum": 1,
   "description": "Maximum PDF pages to render and OCR."
  },
  "pdf_base64": {
   "type": "string",
   "maxLength": 5592406,
   "minLength": 1,
   "description": "One base64-encoded PDF up to 4 MiB decoded; application/pdf data URIs are accepted."
  },
  "file_base64": {
   "type": "string",
   "maxLength": 5592406,
   "minLength": 1,
   "description": "Alias for pdf_base64."
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "psm": 6,
  "text": "--- Page 1 ---\nPrinted text recognized from a scanned PDF",
  "bytes": 48210,
  "chars": 51,
  "model": "tesseract-5+poppler",
  "scope": "pdf_ocr",
  "schema": "delx/ocr/v1",
  "language": "eng",
  "provider": "first-party",
  "truncated": false,
  "confidence": 94.12,
  "attribution": "Text recognized locally by Delx with Tesseract and Poppler; no document is stored, fetched, or sent to a third-party provider.",
  "document_id": "pdf_ocr_6fd2b8d1b8ddf0f3e9a9b2c4d5e6f708",
  "total_pages": 1,
  "input_sha256": "6fd2b8d1b8ddf0f3e9a9b2c4d5e6f7081234567890abcdef1234567890abcdef",
  "output_sha256": "c5f7b2d4c1e7a0f9b6d2e8c3a4f5b6071234567890abcdef1234567890abcdef",
  "pages_observed": 1,
  "words_observed": 9,
  "sale_price_usdc": 0.01,
  "upstream_cost_usd": 0,
  "gross_margin_floor_usd": 0.0095275
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/delx-commerce-pdf-ocr-9bdf08a7/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from commerce.delx.ai](https://www.zero.xyz/host/commerce.delx.ai/llms.txt)
