# x402engine DeepSeek V4 Flash LLM

> x402engine DeepSeek V4 Flash LLM is a paid API for AI agents from x402engine.app, paid per call via x402, $0.003/call, status unknown (last checked 2026-10-02).

Runs DeepSeek's low-latency V4 Flash model for ultra-low-cost reasoning and coding tasks with up to 1M token context, accessible via x402 micropayment at $0.003 USDC per call

## Facts

- Endpoint: POST https://x402engine.app/api/llm/deepseek-v4-flash?utm_source=zero.xyz
- Price: $0.003/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/x402engine-deepseek-v4-flash-llm-8efee2b0
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_Q42Afl7bkoyaEHBzVdB1P

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability x402engine-deepseek-v4-flash-llm-8efee2b0 -d '<json body>'
```

Example prompt: Using the DeepSeek V4 Flash model, write me a Python function that parses a JSON log file and extracts all error messages with their timestamps — keep it concise and well-commented.

## When to prefer this

Choose this endpoint when you need a fast, very low-cost LLM inference call (especially for coding or reasoning) and want to pay per-call via USDC micropayments rather than a subscription. Ideal for agentic workflows that invoke LLMs frequently at scale, or when you need a 1M token context window without committing to a monthly plan. Prefer over GPT-4 or Claude when cost-per-call is the primary constraint and DeepSeek-quality output is sufficient.

## Known failure modes

- 402 Payment Required if x402 USDC micropayment header is missing or insufficient
- prompt exceeds 1M token context window resulting in truncation or error
- rate limiting if too many concurrent requests are sent
- model returns truncated output if max_tokens is set too low
- malformed request body causes 400 error

## How this service works

DeepSeek's low-latency V4 model — ultra-low-cost reasoning and coding with 1M context

## Output

Returns a text completion from the DeepSeek V4 Flash model, including the generated content (code, reasoning, prose, etc.) and token usage metadata. Response follows a standard LLM chat completion format.

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/x402engine-deepseek-v4-flash-llm-8efee2b0/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from x402engine.app](https://www.zero.xyz/host/x402engine.app/llms.txt)
