# ZeroReader Llama 3.2 3B Chat Completion

> ZeroReader Llama 3.2 3B Chat Completion is a paid API for AI agents from api.zeroreader.com, paid per call via x402, $0.002/call, status unknown (last checked 2026-10-02).

Runs chat completions using Meta's Llama 3.2 3B model, offering a fast and cost-effective balance of speed and quality for straightforward text generation tasks.

## Facts

- Endpoint: POST https://api.zeroreader.com/v1/ai/llama-3b?utm_source=zero.xyz
- Price: $0.002/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/zeroreader-llama-3-2-3b-chat-completion-a12ff00c
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_frJ_UC6Xe-f7RmqpCWRWU

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability zeroreader-llama-3-2-3b-chat-completion-a12ff00c -d '<json body>'
```

Example prompt: Send this conversation to the Llama 3.2 3B model with a system message saying 'You are a helpful assistant' and a user message asking 'What are three tips for writing clean Python code?' — keep temperature at 0.7 and max tokens at 512.

## When to prefer this

Choose this endpoint when you need fast, low-cost chat completions for simple or short-form tasks where a 3B parameter model is sufficient — e.g., classification, summarization, Q&A, or lightweight generation. Prefer it over the larger sibling models (17B, 32B, 120B) when latency and cost matter more than peak capability. Avoid it for complex reasoning, math, or code tasks where DeepSeek R1 32B or GPT-OSS 120B would be more appropriate.

## Known failure modes

- Invalid or missing 'messages' array returns a 400 validation error
- Temperature outside [0, 2] range causes a schema validation rejection
- max_tokens exceeding 4096 is rejected with a parameter error
- Payment failure or insufficient USDC balance results in a 402 Payment Required response
- Model unavailability or overload may return a 503 or timeout
- Malformed role values (not 'system', 'user', or 'assistant') cause a 400 error

## How this service works

Llama 3.2 3B — Good balance of speed and quality for simple tasks.

## Output

Returns an OpenAI-compatible chat completion object with a choices array containing the assistant's generated message, a finish_reason ('stop' or 'length'), and a completion ID. The assistant content is a plain text string.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "stream": {
   "type": "boolean",
   "default": false
  },
  "messages": {
   "type": "array",
   "items": {
    "type": "object",
    "required": [
     "role",
     "content"
    ],
    "properties": {
     "role": {
      "enum": [
       "system",
       "user",
       "assistant"
      ],
      "type": "string"
     },
     "content": {
      "type": "string"
     }
    }
   }
  },
  "max_tokens": {
   "type": "integer",
   "default": 1024,
   "maximum": 4096
  },
  "temperature": {
   "type": "number",
   "default": 0.7,
   "maximum": 2,
   "minimum": 0
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "id": "chatcmpl-example",
 "object": "chat.completion",
 "choices": [
  {
   "index": 0,
   "message": {
    "role": "assistant",
    "content": "I'm doing well!"
   },
   "finish_reason": "stop"
  }
 ]
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/zeroreader-llama-3-2-3b-chat-completion-a12ff00c/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from api.zeroreader.com](https://www.zero.xyz/host/api.zeroreader.com/llms.txt)
