# dt0ur.online Image Captioning (ViT-GPT2)

> dt0ur.online Image Captioning (ViT-GPT2) is a paid API for AI agents from dt0ur.online, paid per call via x402, $0.15/call, status unknown (last checked 2026-10-03).

Generates descriptive, context-aware natural language captions from a public image URL using ViT-GPT2 vision-language models.

## Facts

- Endpoint: POST https://dt0ur.online/api/inference/caption?utm_source=zero.xyz
- Price: $0.15/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-03
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/dt0ur-online-image-captioning-vit-gpt2-1250ea2e
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_9xKHBl_t52_TATVeHIKGQ

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability dt0ur-online-image-captioning-vit-gpt2-1250ea2e -d '<json body>'
```

Example prompt: Can you generate a descriptive caption for this image — https://upload.wikimedia.org/wikipedia/commons/thumb/4/47/PNG_transparency_demonstration_1.png/280px-PNG_transparency_demonstration_1.png — using the ViT-GPT2 captioning model?

## When to prefer this

Choose this endpoint when you need fast, automated natural language descriptions of images from public URLs without setting up your own vision model infrastructure. It is well-suited for accessibility alt-text generation, content tagging pipelines, or any workflow where images need to be described in plain English. Prefer this over general-purpose LLM vision calls when you want a dedicated captioning model (ViT-GPT2) optimized specifically for image-to-text description.

## Known failure modes

- Image URL is not publicly accessible or returns a non-200 HTTP status — captioning will fail
- Unsupported image format (e.g. SVG, TIFF, WebP variants) may cause model inference errors
- Very large images may time out or be rejected
- Image contains ambiguous or abstract content — captions may be generic or inaccurate
- Rate limiting or service unavailability if the underlying model host is overloaded

## How this service works

Generates descriptive, context-aware natural language captions from image URLs using ViT-GPT2 vision-language models.

## Output

A natural language string describing the visual content of the image — generated by a ViT-GPT2 vision-language model — capturing objects, scenes, actions, and contextual details visible in the image.

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/dt0ur-online-image-captioning-vit-gpt2-1250ea2e/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from dt0ur.online](https://www.zero.xyz/host/dt0ur.online/llms.txt)
