# x402engine Audio Transcription with Speaker Identification

> x402engine Audio Transcription with Speaker Identification is a paid API for AI agents from x402engine.app, paid per call via x402, $0.1/call, status unknown (last checked 2026-10-02).

Transcribes up to 10 minutes of audio to text with speaker diarization, accepting either a URL or base64-encoded audio

## Facts

- Endpoint: POST https://x402engine.app/api/transcribe?utm_source=zero.xyz
- Price: $0.1/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/x402engine-audio-transcription-with-speaker-identification-3dd2487e
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_2h9e5dPIwPqLA9y8KnLeO

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability x402engine-audio-transcription-with-speaker-identification-3dd2487e -d '<json body>'
```

Example prompt: Can you transcribe this meeting recording for me and tell me who said what? Here's the audio file URL: https://example.com/meeting.mp3

## When to prefer this

Choose this endpoint when you need automatic speech-to-text transcription combined with speaker diarization (who said what) in a single pay-per-call API call, especially for recordings up to 10 minutes. Prefer this over generic transcription APIs when you want speaker labels without setting up your own diarization pipeline, and when paying per use via USDC micropayments (x402 protocol) is acceptable.

## Known failure modes

- Audio file exceeds 10-minute limit — request rejected or truncated
- Invalid or inaccessible audio URL — returns network or fetch error
- Malformed base64 data or missing MIME type prefix — returns parsing error
- Unsupported audio format — returns format error
- Both audio_url and audio_base64 provided simultaneously — schema validation error
- Neither audio_url nor audio_base64 provided — returns missing field error

## How this service works

Convert audio to text with speaker identification — covers up to 10 minutes of audio

## Output

Returns a full text transcript of the audio content, with speaker identification labels (e.g. Speaker 1, Speaker 2) annotating which segments each speaker was responsible for. Covers recordings up to 10 minutes in length.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "$schema": "https://json-schema.org/draft/2020-12/schema",
 "required": [
  "input"
 ],
 "properties": {
  "input": {
   "type": "object",
   "required": [
    "type",
    "bodyType",
    "body",
    "method"
   ],
   "properties": {
    "body": {
     "type": "object",
     "anyOf": [
      {
       "required": [
        "audio_url"
       ]
      },
      {
       "required": [
        "audio_base64"
       ]
      }
     ],
     "properties": {
      "audio_url": {
       "type": "string",
       "example": "https://example.com/audio.mp3",
       "description": "URL of audio file to transcribe (provide audio_url OR audio_base64, not both)"
      },
      "audio_base64": {
       "type": "string",
       "example": "data:audio/mp3;base64,SGVsbG8=",
       "description": "Base64-encoded audio data with MIME type prefix (provide audio_url OR audio_base64, not both)"
      },
      "audio_mimetype": {
       "type": "string",
       "example": "audio/mp3",
       "description": "MIME type when using audio_base64 (e.g. audio/mp3, audio/wav)"
      }
     },
     "additionalProperties": false
    },
    "type": {
     "type": "string",
     "const": "http"
    },
    "method": {
     "enum": [
      "POST"
     ],
     "type": "string"
    },
    "bodyType": {
     "enum": [
      "json",
      "form-data",
      "text"
     ],
     "type": "string"
    }
   },
   "additionalProperties": false
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/x402engine-audio-transcription-with-speaker-identification-3dd2487e/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from x402engine.app](https://www.zero.xyz/host/x402engine.app/llms.txt)
