# agent402.tools Pro Chat Completions

> agent402.tools Pro Chat Completions is a paid API for AI agents from agent402.tools, paid per call via x402, $0.1/call, status unknown (last checked 2026-09-13).

OpenAI-compatible chat completions at pro tier supporting GPT-4o, GPT-4.1, Claude Sonnet, Gemini Pro, and Grok with higher input/output caps, paid per call in USDC via x402

## Facts

- Endpoint: POST https://agent402.tools/v1/pro/chat/completions
- Price: $0.1/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-13
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/agent402-tools-pro-chat-completions-175d6487
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_TlvuIHFyJBaKGPZ6tmQi7

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability agent402-tools-pro-chat-completions-175d6487 -d '<json body>'
```

Example prompt: Send this 40,000-character research document to Claude Sonnet via the pro chat completions endpoint and ask it to summarize the key findings — pay per call in USDC, and make sure zero-data-retention mode is on.

## When to prefer this

Use this endpoint when you need access to top-tier models (GPT-4o, GPT-4.1, Claude Sonnet, Gemini Pro, Grok) via a single OpenAI-compatible interface, especially when paying per call in USDC is preferred over a subscription. Ideal for agents that need large context windows (up to 48k chars input), want to avoid managing multiple API keys, require zero-data-retention routing for sensitive workloads, or want crypto-native pay-as-you-go LLM inference. Choose over the standard tier when you need the higher-capability models or larger I/O caps.

## Known failure modes

- Payment failure: x402 payment not accepted or wallet has insufficient USDC balance
- Model not available: requested model not on the pro-tier allowlist (check GET /v1/models)
- ZDR routing failure: zdr=true requested but no zero-data-retention provider available for that model, walks failover chain and may error
- Input too large: input exceeds 48k character cap
- Max tokens exceeded: max_tokens clamped to tier maximum (4096)
- Rate limiting or upstream provider outage causing 5xx errors
- Invalid message format: messages array malformed or unsupported content block type

## How this service works

OpenAI-compatible chat completions, pro tier: gpt-4o, gpt-4.1, claude sonnet, gemini pro, grok - paid per call in USDC over x402. Same wire format as /v1/chat/completions with higher input/output caps (48k chars in, 4096 tokens out).

## Output

An OpenAI-compatible chat completion response object containing the assistant's generated message, finish reason, and token usage stats. Output is capped at 4096 tokens with up to 48k characters of input accepted. The response follows the same wire format as /v1/chat/completions so it works with any OpenAI SDK.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "zdr": {
   "type": "boolean",
   "description": "Optional - true routes only to zero-data-retention providers (OpenRouter provider.zdr); the only provider preference a caller may set."
  },
  "model": {
   "type": "string",
   "description": "Model id - OpenRouter form (openai/gpt-4o-mini) or bare OpenAI form (gpt-4o-mini). GET /v1/models lists the allowlist per tier. Optional: omit it and the tier serves its documented default (x402.defaultModel on /v1/models), named back in agent402_default_model; the price does not change"
  },
  "tools": {
   "type": "array",
   "description": "Optional - OpenAI function tools {type:\"function\", function:{...}}, or a tool namespace {type:\"namespace\", name, tools:[...]} (flattened into its functions). The pro and premium routes also accept the bounded server tools openrouter:web_search, openrouter:web_fetch and openrouter:datetime with server-owned limits (GET /v1/models lists them); stop_server_tools_when and max_tool_calls are refused. A request with a server tool is never served from the prompt cache."
  },
  "messages": {
   "type": "array",
   "description": "OpenAI chat messages: [{role, content}] - text and image_url content blocks supported"
  },
  "reasoning": {
   "type": "object",
   "description": "Optional - {effort: \"none\"|\"minimal\"|\"low\"|\"medium\"|\"high\"|\"xhigh\"|\"max\", max_tokens?, exclude?, enabled?}. Reasoning tokens count against max_tokens. Omitted: low effort on the budget tiers, the model default on premium. reasoning_effort (string) is accepted as an alias."
  },
  "max_tokens": {
   "type": "number",
   "description": "Output token cap (clamped to the tier maximum)"
  },
  "cache_control": {
   "description": "Optional - prompt caching preference. Default ON ({type:\"ephemeral\"}, 5-minute TTL): repeated prefixes across your turns are served from the provider cache (same price to you). Send false to disable. ttl:\"1h\" is not offered."
  },
  "max_completion_tokens": {
   "type": "integer",
   "description": "Optional - alias of max_tokens (newer OpenAI SDKs send this)."
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "id": "gen-…",
  "model": "openai/gpt-4o",
  "usage": {
   "total_tokens": 13,
   "prompt_tokens": 12,
   "completion_tokens": 1
  },
  "object": "chat.completion",
  "choices": [
   {
    "index": 0,
    "message": {
     "role": "assistant",
     "content": "OK"
    },
    "finish_reason": "stop"
   }
  ],
  "created": 1750000000
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/agent402-tools-pro-chat-completions-175d6487/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from agent402.tools](https://www.zero.xyz/host/agent402.tools/llms.txt)
