# SiteCheck API – Speech-to-Text Transcription

> SiteCheck API – Speech-to-Text Transcription is a paid API for AI agents from api.sitecheck-api.workers.dev, paid per call via x402, $0.01/call, status unknown (last checked 2026-10-02).

Transcribes audio files (via URL or base64) into text with timestamps, language detection, and segment-level detail using Whisper Large v3 Turbo

## Facts

- Endpoint: POST https://api.sitecheck-api.workers.dev/api/transcribe?utm_source=zero.xyz
- Price: $0.01/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/sitecheck-api-speech-to-text-transcription-add41a34
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_6fJ9J0UVP1oF5z7z2JELk

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability sitecheck-api-speech-to-text-transcription-add41a34 -d '<json body>'
```

Example prompt: Can you transcribe this audio file for me? Here's the URL: https://example.com/interview.mp3 — I think it's in English but you can auto-detect the language.

## When to prefer this

Choose this endpoint when you need a pay-per-call, keyless speech-to-text transcription with no signup overhead, especially when paying via USDC on Base via x402. It returns rich segment-level timestamps alongside full transcripts, making it ideal for captioning or meeting notes. Prefer it over managed STT services when you want crypto-native micropayments and no API key management.

## Known failure modes

- Audio URL is inaccessible or returns a non-200 response — transcription fails
- Unsupported audio format provided — error or empty result
- Base64 audio is malformed or too large — request rejected
- Language code is invalid — may fall back to auto-detect or error
- Very long audio files may time out on the Cloudflare Workers runtime
- Payment not included or insufficient USDC — 402 Payment Required response

## How this service works

Speech-to-text with Whisper large-v3-turbo: send an https audio URL (or base64) up to 25 MB, get the transcript, language and timestamped segments.

## Output

Returns a JSON object with the full transcribed text, the Whisper model used, total audio duration in seconds, the detected or specified language code, and an array of timed segments each containing start time, end time, and the spoken text in that segment.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "url": {
   "type": "string",
   "description": "https URL of an audio file (mp3, wav, m4a...)"
  },
  "language": {
   "type": "string",
   "description": "Optional ISO code, e.g. en"
  },
  "audio_base64": {
   "type": "string",
   "description": "Alternative to url"
  }
 }
}
```

## Response schema (JSON Schema)

```json
{
 "type": "json",
 "example": {
  "text": "Hello and welcome...",
  "model": "@cf/openai/whisper-large-v3-turbo",
  "duration": 42.1,
  "language": "en",
  "segments": [
   {
    "end": 3.2,
    "text": "Hello and welcome",
    "start": 0
   }
  ]
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/sitecheck-api-speech-to-text-transcription-add41a34/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from api.sitecheck-api.workers.dev](https://www.zero.xyz/host/api.sitecheck-api.workers.dev/llms.txt)
