# ZeroReader Llama 3.3 70B (Fast) Chat Completion

> ZeroReader Llama 3.3 70B (Fast) Chat Completion is a paid API for AI agents from api.zeroreader.com, paid per call via x402, $0.008/call, status unknown (last checked 2026-09-30).

Runs inference against Meta's Llama 3.3 70B flagship open-source LLM via Cloudflare Workers AI, returning a chat completion response in OpenAI-compatible format

## Facts

- Endpoint: POST https://api.zeroreader.com/v1/ai/llama-70b?utm_source=zero.xyz
- Price: $0.008/call
- Payment: x402
- Status: unknown
- Last checked: 2026-09-30
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/zeroreader-llama-3-3-70b-fast-chat-completion-0a092c4c
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_jVMJi0WRdzJ2P61TJUbrP

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability zeroreader-llama-3-3-70b-fast-chat-completion-0a092c4c -d '<json body>'
```

Example prompt: Send this conversation to Llama 3.3 70B and get a response — system prompt: 'You are a helpful assistant', user message: 'Explain the difference between supervised and unsupervised learning in plain English', with temperature 0.7 and up to 1024 tokens.

## When to prefer this

Choose this endpoint when you need a high-quality, flagship open-source LLM (Llama 3.3 70B) with fast inference, pay-per-call USDC micropayment pricing (no subscription required), and OpenAI-compatible response format. Prefer this over smaller siblings (3B, 7B) when task complexity demands best-in-class open-source quality, and over reasoning-specialized siblings (DeepSeek R1) when you need general-purpose chat rather than chain-of-thought math/logic.

## Known failure modes

- Payment not provided or invalid x402 payment header — 402 Payment Required
- Temperature out of range (>2 or <0) — validation error
- max_tokens exceeds 4096 — validation error
- Messages array missing required role or content fields — 400 Bad Request
- Model overloaded or Cloudflare Workers AI backend unavailable — 503 Service Unavailable
- Malformed JSON body — 400 Bad Request

## How this service works

Llama 3.3 70B (Fast) — Flagship open-source LLM. Best quality on CF Workers AI.

## Output

An OpenAI-compatible chat completion object containing a unique completion ID, the assistant's generated text in the message content field, a finish_reason (e.g. 'stop'), and a choice index. The response is a single JSON object (non-streaming by default).

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "stream": {
   "type": "boolean",
   "default": false
  },
  "messages": {
   "type": "array",
   "items": {
    "type": "object",
    "required": [
     "role",
     "content"
    ],
    "properties": {
     "role": {
      "enum": [
       "system",
       "user",
       "assistant"
      ],
      "type": "string"
     },
     "content": {
      "type": "string"
     }
    }
   }
  },
  "max_tokens": {
   "type": "integer",
   "default": 1024,
   "maximum": 4096
  },
  "temperature": {
   "type": "number",
   "default": 0.7,
   "maximum": 2,
   "minimum": 0
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "id": "chatcmpl-example",
 "object": "chat.completion",
 "choices": [
  {
   "index": 0,
   "message": {
    "role": "assistant",
    "content": "I'm doing well!"
   },
   "finish_reason": "stop"
  }
 ]
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/zeroreader-llama-3-3-70b-fast-chat-completion-0a092c4c/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from api.zeroreader.com](https://www.zero.xyz/host/api.zeroreader.com/llms.txt)
