# Agent402 Premium AI Messages (x402 Pay-Per-Call LLM Inference)

> Agent402 Premium AI Messages (x402 Pay-Per-Call LLM Inference) is a paid API for AI agents from agent402.tools, paid per call via x402, $0.5/call, status unknown (last checked 2026-09-15).

Proxies Anthropic-compatible chat completion requests to frontier LLMs, billed per call in USDC via x402 with no API key required

## Facts

- Endpoint: POST https://agent402.tools/v1/premium/messages
- Price: $0.5/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-15
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/agent402-premium-ai-messages-x402-pay-per-call-llm-inference-ec95ce7a
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_5NTBhuDOkaoAa2_LWbLVX

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability agent402-premium-ai-messages-x402-pay-per-call-llm-inference-ec95ce7a -d '<json body>'
```

Example prompt: Ask Claude Sonnet 5 via x402 pay-per-call to explain how proof-of-work consensus works — use the auto model tier, allow up to 1024 output tokens, and stream the response.

## When to prefer this

Choose this endpoint when you need LLM inference (especially Claude-family models) without managing API keys or subscriptions — ideal for AI agents that self-fund via a USDC wallet, scenarios requiring anonymous or pseudonymous access (wallet-as-identity), compliance use cases needing zero-data-retention, or when building on the x402 / MPP payment standard. Prefer alternatives (direct Anthropic API, OpenAI) if you already have managed credentials, need non-Claude models not on the allowlist, or cannot handle x402 payment flows.

## Known failure modes

- 402 Payment Required if wallet has insufficient USDC balance
- Model ID not allowlisted for the caller's tier returns an error
- max_tokens exceeds tier output cap and gets clamped or rejected
- Malformed messages array (wrong role sequence) causes 400 bad request
- Zero-data-retention requested but no ZDR-compliant provider available for the selected model
- Streaming connection dropped mid-response leaves partial SSE stream
- Tool schema validation failure if client-supplied tools have invalid input_schema

## How this service works

Anthropic Messages API over x402 - point the Anthropic SDK (or Claude Code / the Agent SDK) at base_url https://agent402.tools/v1/premium and pay $0.50 per call in USDC, no API key, no signup. Same models, caps and price as this tier's /chat/completions route; any model here is served through the Messages wire (Claude natively, others translated). Up to 200,000 input chars and 8192 output tokens; streaming supported.

## Output

Returns an Anthropic-format message object containing the assistant role, one or more content blocks (text and/or tool_use), the stop_reason (end_turn, tool_use, max_tokens, etc.), the resolved model name, and input/output token counts. When streaming is enabled, delivers Anthropic SSE events (message_start through message_stop).

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "zdr": {
   "type": "boolean",
   "description": "Optional - zero-data-retention providers only"
  },
  "model": {
   "type": "string",
   "description": "Model id (OpenRouter naming, e.g. anthropic/claude-sonnet-5) - allowlisted per tier; omit (or \"auto\") on the auto tier"
  },
  "tools": {
   "type": "array",
   "description": "Optional client tools {name, description, input_schema}; server/built-in tools are not served"
  },
  "stream": {
   "type": "boolean",
   "description": "Anthropic SSE (message_start … message_stop)"
  },
  "system": {
   "type": "string",
   "description": "Optional system prompt (string or text blocks)"
  },
  "messages": {
   "type": "array",
   "description": "Anthropic messages: {role: user|assistant, content: string | [text|image|tool_use|tool_result blocks]}"
  },
  "thinking": {
   "type": "object",
   "description": "Optional {type:\"enabled\", budget_tokens} | {type:\"adaptive\"} | {type:\"disabled\"} - thinking tokens are output tokens"
  },
  "max_tokens": {
   "type": "integer",
   "description": "Required by the Messages API; clamped to the tier's output cap"
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "id": "msg_…",
  "role": "assistant",
  "type": "message",
  "model": "anthropic/claude-sonnet-5",
  "usage": {
   "input_tokens": 14,
   "output_tokens": 18
  },
  "content": [
   {
    "text": "x402 is an HTTP-native way for agents to pay per request with USDC.",
    "type": "text"
   }
  ],
  "stop_reason": "end_turn"
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/agent402-premium-ai-messages-x402-pay-per-call-llm-inference-ec95ce7a/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from agent402.tools](https://www.zero.xyz/host/agent402.tools/llms.txt)
