# Aayat AI Fast LLM Chat (Llama 3.2 3B)

> Aayat AI Fast LLM Chat (Llama 3.2 3B) is a paid API for AI agents from aayatai.com, paid per call via x402, $0.003/call, status unknown (last checked 2026-10-02).

Pay-per-call LLM inference using Llama 3.2 3B for fast, cheap text classification, extraction, short-answer generation, and routing tasks — no API key required.

## Facts

- Endpoint: POST https://aayatai.com/chat/fast?utm_source=zero.xyz
- Price: $0.003/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/aayat-ai-fast-llm-chat-llama-3-2-3b-373558df
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_SPWEBVCCiA9-U9E4PmqY5

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability aayat-ai-fast-llm-chat-llama-3-2-3b-373558df -d '<json body>'
```

Example prompt: Classify this customer message as 'complaint', 'question', or 'compliment' and return only valid JSON with a 'label' field: 'My order arrived two days late and the packaging was damaged.'

## When to prefer this

Choose this endpoint when you need fast, cheap LLM inference for simple tasks like classification, short extraction, or routing and want to avoid API key provisioning or account setup. It is ideal for high-volume, low-complexity agentic subtasks where cost per call matters and a 3B-parameter model is sufficient. Prefer larger hosted LLMs (GPT-4, Claude) when the task requires deep reasoning, long context, or high accuracy on complex language understanding.

## Known failure modes

- Prompt or messages exceeding ~15,000 characters may be rejected or truncated
- max_tokens above 1024 will be rejected by schema validation
- Malformed request body (missing both prompt and messages) returns an error and is not charged
- Ambiguous or very long prompts may produce truncated or low-quality outputs due to the small 3B model size
- Network timeouts on the caller side; failed calls are not billed

## How this service works

Pay-per-call LLM chat (Llama 3.2 3B): Fast, cheap model for classification, extraction, short answers and routing. No API key or account. POST JSON {"prompt": "..."} or OpenAI-style {"messages": [{"role": "user", "content": "..."}]}, optional "system", "max_tokens" (up to 1024), "temperature", "json": true. Up to about 15,000 characters of English in. Failed calls are not charged.

## Output

Returns a JSON object containing the generated text under 'text', the model identifier under 'model' (e.g. '@cf/meta/llama-3.2-3b-instruct'), and a 'usage' object with 'promptTokens' and 'completionTokens' counts. When json mode is enabled, the 'text' field contains valid JSON output.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "json": {
   "type": "boolean",
   "default": false,
   "description": "Ask for JSON-only output."
  },
  "prompt": {
   "type": "string",
   "maxLength": 20000,
   "description": "A single user message (use this or messages)."
  },
  "system": {
   "type": "string",
   "maxLength": 8000,
   "description": "System instructions."
  },
  "messages": {
   "type": "array",
   "items": {
    "type": "object",
    "properties": {
     "role": {
      "enum": [
       "system",
       "user",
       "assistant"
      ],
      "type": "string"
     },
     "content": {
      "type": "string"
     }
    }
   },
   "maxItems": 50,
   "description": "Conversation so far, OpenAI style: [{role, content}]."
  },
  "max_tokens": {
   "type": "integer",
   "maximum": 1024,
   "minimum": 1,
   "description": "Most tokens to generate (default 512)."
  },
  "temperature": {
   "type": "number",
   "default": 0.3,
   "maximum": 2,
   "minimum": 0,
   "description": "Randomness, 0-2."
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "text": "x402 is an open protocol that lets clients pay for HTTP requests with stablecoins using the 402 status code.",
  "model": "@cf/meta/llama-3.2-3b-instruct",
  "usage": {
   "promptTokens": 24,
   "completionTokens": 26
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/aayat-ai-fast-llm-chat-llama-3-2-3b-373558df/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from aayatai.com](https://www.zero.xyz/host/aayatai.com/llms.txt)
