# Xona Agent LLM Inference (NVIDIA NIM / OpenRouter Fallback)

> Xona Agent LLM Inference (NVIDIA NIM / OpenRouter Fallback) is a paid API for AI agents from api.xona-agent.com, paid per call via x402, $0.015/call, status unknown (last checked 2026-10-02).

OpenAI-compatible chat completions with automatic provider fallback (NVIDIA NIM → OpenRouter), supporting frontier models like DeepSeek, Kimi, GLM, and Nemotron via pay-per-call x402 pricing.

## Facts

- Endpoint: POST https://api.xona-agent.com/llm/nim?utm_source=zero.xyz
- Price: $0.015/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/xona-agent-llm-inference-nvidia-nim-openrouter-fallback-48c49d35
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_bBPSwl_H868XVOOnggFqA

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability xona-agent-llm-inference-nvidia-nim-openrouter-fallback-48c49d35 -d '<json body>'
```

Example prompt: Using the deepseek-v4-pro model, answer this question with up to 512 output tokens: 'Explain the difference between gradient descent and stochastic gradient descent in simple terms.'

## When to prefer this

Prefer this endpoint when you need frontier model inference (DeepSeek-v4-pro, Nemotron-3-ultra-550b, Kimi-k2.6, GLM-5.1) with high reliability via automatic provider fallback, and want pay-per-call x402 micropayment pricing instead of a subscription. Ideal for agents that need sporadic, usage-based LLM calls without managing separate provider API keys.

## Known failure modes

- Model not available — if NVIDIA NIM is unavailable and OpenRouter fallback also fails, returns an error
- max_tokens exceeded — request rejected if max_tokens > 8192
- Invalid model ID — returns error if model field does not match one of the four supported model IDs
- Payment failure — x402 charge not resolved, request blocked
- Rate limiting — too many concurrent requests may be throttled

## How this service works

Frontier-model inference with automatic provider fallback (NVIDIA NIM → OpenRouter) for reliability. OpenAI-compatible chat completions; select a model via the body `model` field. Usage-based pricing (input + output tokens at provider cost + 10%); the x402 per-call charge is quoted up front as an upper bound, so set max_tokens to lower it. Also available drop-in at POST /v1/chat/completions (API key + credits, billed exactly per token).

## Output

An OpenAI-compatible chat completion response containing the assistant's generated text, along with token usage details (input + output tokens). The response follows the standard ChatCompletion object format.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "model": {
   "type": "string",
   "default": "deepseek-v4-pro",
   "description": "Model id (usage-based price: input + output tokens at provider cost + 10%). One of: kimi-k2.6, glm-5.1, nemotron-3-ultra-550b, deepseek-v4-pro"
  },
  "prompt": {
   "type": "string",
   "description": "The user prompt (or use messages)"
  },
  "messages": {
   "type": "array",
   "items": {
    "type": "string"
   },
   "description": "Optional OpenAI-style messages array (overrides prompt/system_prompt)"
  },
  "max_tokens": {
   "type": "number",
   "description": "Cap on output tokens (max 8192). Lowers the up-front per-call charge — unset quotes the full cap as an upper bound."
  },
  "temperature": {
   "type": "number",
   "description": "Optional sampling temperature"
  },
  "system_prompt": {
   "type": "string",
   "description": "Optional system prompt"
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/xona-agent-llm-inference-nvidia-nim-openrouter-fallback-48c49d35/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from api.xona-agent.com](https://www.zero.xyz/host/api.xona-agent.com/llms.txt)
