# Llama API (Pay-Per-Call Chat Completions)

> Llama API (Pay-Per-Call Chat Completions) is a paid API for AI agents from x402.agentindex.world, paid per call via x402, $0.001/call, status unknown (last checked 2026-10-01).

Pay-per-call OpenAI-format chat completions powered by Meta Llama 4 Maverick via DeepInfra, with no API key required and fixed per-call pricing.

## Facts

- Endpoint: GET https://x402.agentindex.world/llm/llama?utm_source=zero.xyz
- Price: $0.001/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-01
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/llama-api-pay-per-call-chat-completions-dd28e0f4
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_VuvUH8GBBtvzvCpw1Zyvf

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability llama-api-pay-per-call-chat-completions-dd28e0f4
```

Example prompt: Ask Llama 4 Maverick: 'Explain the difference between supervised and unsupervised learning in 3 short paragraphs' — use up to 512 tokens for the reply.

## When to prefer this

Choose this endpoint when you need a pay-per-call LLM completion with no API key setup, no subscription, and predictable per-call USDC pricing. It is ideal for AI agents that need sporadic or burst LLM calls without committing to a provider account, or for prototyping pipelines using the OpenAI message format against a capable open-weights model (Llama 4 Maverick). Prefer alternatives if you need streaming responses, model selection flexibility, refunds for unused tokens, or latency under a few seconds consistently.

## Known failure modes

- 20-second server-side timeout fires — request is not charged but no completion is returned
- max_tokens exceeds 4096 cap — request may be rejected or capped
- messages array is empty or malformed — 400-level error returned
- USDC payment fails or is insufficient — payment-gated 402 response, no completion returned
- Client timeout shorter than 30s causes premature disconnect before server responds

## How this service works

Llama API - pay per call, no API key. OpenAI-format chat completions on meta-llama/llama-4-maverick, pinned to DeepInfra for reliability. You sign a fixed price computed from your max_tokens; unused tokens are not refunded. 20s server-side timeout - never charged if it fires; set your client timeout to 30s. Try GET /llm/llama/sample.

## Output

An OpenAI-compatible chat completion response containing the assistant's generated text, produced by Meta Llama 4 Maverick. The response reflects the conversation history provided in the messages array and is bounded by the max_tokens parameter. Unused tokens within the ceiling are not refunded. The server enforces a 20-second timeout and charges are not applied if the timeout fires.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "messages": {
   "type": "array",
   "items": {
    "type": "object"
   },
   "minItems": 1,
   "description": "OpenAI-format chat messages: [{role, content}, ...]."
  },
  "max_tokens": {
   "type": "integer",
   "description": "Max completion tokens, capped at 4096 - also bounds the price ceiling."
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/llama-api-pay-per-call-chat-completions-dd28e0f4/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from x402.agentindex.world](https://www.zero.xyz/host/x402.agentindex.world/llms.txt)
