# Agent402 Auto Responses - Pay-Per-Call AI Inference Gateway

> Agent402 Auto Responses - Pay-Per-Call AI Inference Gateway is a paid API for AI agents from agent402.tools, paid per call via x402, $0.01/call, status unknown (last checked 2026-09-14).

Routes a natural-language or multi-turn input to an AI language model and returns a completed response, billed at $0.01 USDC per call with no API key required — wallet serves as identity via x402.

## Facts

- Endpoint: POST https://agent402.tools/v1/auto/responses
- Price: $0.01/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-14
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/agent402-auto-responses-pay-per-call-ai-inference-gateway-e4296557
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_V8h97UFx7aZFUze7Zgun_

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability agent402-auto-responses-pay-per-call-ai-inference-gateway-e4296557 -d '<json body>'
```

Example prompt: Ask the auto-tier model to explain what the x402 protocol is in one sentence, and return the answer as plain text — charge my USDC wallet per call.

## When to prefer this

Choose this endpoint when you need anonymous, keyless LLM inference billed per call in USDC, especially inside an agentic workflow that already uses the x402 payment protocol. It is ideal when you want automatic model routing without managing multiple provider API keys, or when your agent needs to pay for intelligence on demand from a crypto wallet. Prefer it over provider-direct APIs when you want cost flexibility, no subscription lock-in, and access to 500+ tools in the same ecosystem.

## Known failure modes

- Insufficient USDC balance in wallet — payment rejected before inference runs
- Model ID not on the allowlist for the caller's tier — returns an error listing valid models
- Input tokens exceed the tier's context cap — request is rejected or truncated
- Malformed input array (missing role/content structure) — schema validation error
- zero-data-retention flag set but no ZDR-compliant provider available for the requested model
- Streaming disconnection mid-response — partial SSE stream with no completed event
- Rate limits on the underlying OpenRouter provider — upstream 429 propagated back

## How this service works

OpenAI Responses API over x402 - point the OpenAI SDK's responses.create() (or the OpenAI Agents SDK) at base_url https://agent402.tools/v1/auto and pay $0.01 per call in USDC, no API key, no signup. Same models, caps and price as this tier's /chat/completions route; any model here is served through the Responses wire.

## Output

A JSON object conforming to the OpenAI Responses API shape: includes a response ID, the model used, token usage breakdown (input/output/total), a status field, and an output array containing the assistant message with text content and any annotations. If streaming is enabled, Server-Sent Events are returned instead.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "zdr": {
   "type": "boolean",
   "description": "Optional - zero-data-retention providers only"
  },
  "text": {
   "type": "object",
   "description": "Optional {format: {type: \"text\"|\"json_schema\"|\"json_object\", ...}}"
  },
  "input": {
   "description": "A string, or an array of input items ({role, content} messages with input_text / input_image parts, function_call, function_call_output)"
  },
  "model": {
   "type": "string",
   "description": "Model id (OpenRouter naming) - allowlisted per tier; omit (or \"auto\") on the auto tier"
  },
  "tools": {
   "type": "array",
   "description": "Optional function tools ({type:\"function\", name, parameters}); server-side tools are not served"
  },
  "stream": {
   "type": "boolean",
   "description": "Responses SSE events (response.created … response.completed)"
  },
  "reasoning": {
   "type": "object",
   "description": "Optional {effort: \"none\"|\"minimal\"|\"low\"|\"medium\"|\"high\"|\"xhigh\"|\"max\"} - reasoning tokens count against max_output_tokens"
  },
  "instructions": {
   "type": "string",
   "description": "Optional system/developer instructions"
  },
  "max_output_tokens": {
   "type": "integer",
   "description": "Optional output cap (clamped to the tier cap)"
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "id": "resp_…",
  "model": "openai/gpt-4o-mini",
  "usage": {
   "input_tokens": 14,
   "total_tokens": 32,
   "output_tokens": 18
  },
  "object": "response",
  "output": [
   {
    "id": "msg_…",
    "role": "assistant",
    "type": "message",
    "status": "completed",
    "content": [
     {
      "text": "x402 is an HTTP-native way for agents to pay per request with USDC.",
      "type": "output_text",
      "annotations": []
     }
    ]
   }
  ],
  "status": "completed"
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/agent402-auto-responses-pay-per-call-ai-inference-gateway-e4296557/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from agent402.tools](https://www.zero.xyz/host/agent402.tools/llms.txt)
