# Agent402 AI Messages API — Pay-Per-Call LLM Inference

> Agent402 AI Messages API — Pay-Per-Call LLM Inference is a paid API for AI agents from agent402.tools, paid per call via x402, $0.02/call, status unknown (last checked 2026-09-15).

Sends messages to a large language model via the Anthropic Messages API format, paying $0.02 USDC per call over x402 with no API keys required

## Facts

- Endpoint: POST https://agent402.tools/v1/messages
- Price: $0.02/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-15
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/agent402-ai-messages-api-pay-per-call-llm-inference-0d8feed3
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_nxLrEjRV8SdPI7LKhMWus

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability agent402-ai-messages-api-pay-per-call-llm-inference-0d8feed3 -d '<json body>'
```

Example prompt: Ask Claude Sonnet (anthropic/claude-sonnet-5) to explain what x402 is in one sentence — pay per call with my USDC wallet, no API key needed, and stream the response back.

## When to prefer this

Choose this endpoint when you need LLM inference with no signup friction, no API key management, and wallet-based micropayment identity — ideal for autonomous AI agents that self-fund calls via x402, privacy-sensitive workloads needing zero-data-retention providers, or multi-agent pipelines that need per-call cost accounting in USDC. Prefer it over direct Anthropic or OpenRouter access when the caller is an on-chain agent, when you want pay-as-you-go without monthly subscriptions, or when you need access to the open x402 ecosystem's smart order routing across 558+ tools.

## Known failure modes

- Payment not received or insufficient USDC balance — 402 Payment Required returned before inference runs
- Requested model not allowlisted for caller's tier — model ID rejected with error
- max_tokens exceeds tier output cap — clamped silently or rejected
- Malformed messages array (missing role or content) — 400 validation error
- ZDR flag set but no ZDR-compliant provider available — fallback or error
- Stream connection dropped mid-response — partial SSE events received
- Rate limit hit at the routing layer — 429 Too Many Requests

## How this service works

Anthropic Messages API over x402 - point the Anthropic SDK (or Claude Code / the Agent SDK) at base_url https://agent402.tools/v1 and pay $0.02 per call in USDC, no API key, no signup. Same models, caps and price as this tier's /chat/completions route; any model here is served through the Messages wire (Claude natively, others translated). Up to 32,000 input chars and 2048 output tokens; streaming supported.

## Output

Returns an Anthropic-format message object with an assistant role, content array of text or tool_use blocks, stop_reason, model identifier, and input/output token counts — identical to the Anthropic Messages API response envelope.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "zdr": {
   "type": "boolean",
   "description": "Optional - zero-data-retention providers only"
  },
  "model": {
   "type": "string",
   "description": "Model id (OpenRouter naming, e.g. anthropic/claude-sonnet-5) - allowlisted per tier; omit (or \"auto\") on the auto tier"
  },
  "tools": {
   "type": "array",
   "description": "Optional client tools {name, description, input_schema}; server/built-in tools are not served"
  },
  "stream": {
   "type": "boolean",
   "description": "Anthropic SSE (message_start … message_stop)"
  },
  "system": {
   "type": "string",
   "description": "Optional system prompt (string or text blocks)"
  },
  "messages": {
   "type": "array",
   "description": "Anthropic messages: {role: user|assistant, content: string | [text|image|tool_use|tool_result blocks]}"
  },
  "thinking": {
   "type": "object",
   "description": "Optional {type:\"enabled\", budget_tokens} | {type:\"adaptive\"} | {type:\"disabled\"} - thinking tokens are output tokens"
  },
  "max_tokens": {
   "type": "integer",
   "description": "Required by the Messages API; clamped to the tier's output cap"
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "id": "msg_…",
  "role": "assistant",
  "type": "message",
  "model": "anthropic/claude-sonnet-5",
  "usage": {
   "input_tokens": 14,
   "output_tokens": 18
  },
  "content": [
   {
    "text": "x402 is an HTTP-native way for agents to pay per request with USDC.",
    "type": "text"
   }
  ],
  "stop_reason": "end_turn"
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/agent402-ai-messages-api-pay-per-call-llm-inference-0d8feed3/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from agent402.tools](https://www.zero.xyz/host/agent402.tools/llms.txt)
