# dt0ur.online Speech Transcription

> dt0ur.online Speech Transcription is a paid API for AI agents from dt0ur.online, paid per call via x402, $0.2/call, status unknown (last checked 2026-10-02).

Transcribes spoken audio from a WAV file URL into timestamped English text using Whisper neural speech recognition.

## Facts

- Endpoint: POST https://dt0ur.online/api/speech/transcribe?utm_source=zero.xyz
- Price: $0.2/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/dt0ur-online-speech-transcription-4217b559
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_kndfGJjKJ8MHPbkRjeYlr

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability dt0ur-online-speech-transcription-4217b559 -d '<json body>'
```

Example prompt: Can you transcribe the spoken audio in this WAV file for me? Here's the public URL: https://mybucket.s3.amazonaws.com/interview.wav — I need the English text with timestamps.

## When to prefer this

Choose this endpoint when you need to convert a WAV audio file hosted at a public URL into English text with timestamps, especially in automated pipelines where Whisper-quality transcription is required. Prefer it for single-file transcription tasks where the audio is already in WAV/PCM format. Avoid for non-WAV formats.

## Known failure modes

- Non-WAV formats (MP3, OGG, FLAC) cause 500 errors — only PCM WAV is supported
- Inaccessible or private audio URLs result in failure to fetch the file
- Very long audio files may time out or exceed processing limits
- Audio with heavy background noise or non-English speech may produce inaccurate transcriptions
- Malformed or corrupt WAV files cause decoding errors

## How this service works

Transcribes spoken audio streams into timestamped English text using embedded Whisper neural speech recognition models.

## Output

Returns a timestamped English-language transcript of the spoken content in the provided WAV audio file, with text segments aligned to time positions in the audio.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "audio_url": {
   "type": "string",
   "description": "Public URL or file:// URI to a WAV (PCM) audio file. Only WAV is decoded; other formats (OGG, MP3, ...) fail with a 500."
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/dt0ur-online-speech-transcription-4217b559/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from dt0ur.online](https://www.zero.xyz/host/dt0ur.online/llms.txt)
