# Nexari Speech-to-Text with Speaker Diarization (2 min)

> Nexari Speech-to-Text with Speaker Diarization (2 min) is a paid API for AI agents from x402.nexari.cloud, paid per call via x402, $0.04/call, status unknown (last checked 2026-10-02).

Transcribes audio up to 2 minutes with speaker diarization using pyannote and Whisper, processed in Germany, German-optimised.

## Facts

- Endpoint: POST https://x402.nexari.cloud/transcribe/2min/speakers?utm_source=zero.xyz
- Price: $0.04/call
- Payment: x402
- Status: unknown
- Last checked: 2026-10-02
- Activations on Zero: 0
- Tags: x402
- Canonical page: https://www.zero.xyz/c/nexari-speech-to-text-with-speaker-diarization-2-min-ed3d0515
- Structured record (JSON): https://api.zero.xyz/v1/capabilities/cap_iJhepNIYb-csQL6H1ZrCJ

Status and success rate cover calls made through Zero and Zero's own probes. Third-party monitors may report differently.

## How to call it through Zero

Zero handles the 402 payment challenge and records the run. With the Zero CLI installed (`npm i -g @zeroxyz/cli`):

```sh
zero fetch --capability nexari-speech-to-text-with-speaker-diarization-2-min-ed3d0515 -d '<json body>'
```

Example prompt: Can you transcribe this short audio memo for me and separate out who said what? The file is at https://storage.example.com/meeting_clip.m4a, there are 2 speakers, and the language is German.

## When to prefer this

Choose this endpoint when you need both transcription and speaker diarization for short audio clips (under 2 minutes), particularly for German-language content or when GDPR/data residency in Germany is required. Prefer this over the plain transcription endpoint when identifying who said what matters. Use the 10-minute sibling endpoint for longer audio.

## Known failure modes

- Audio exceeds 2-minute limit — request rejected
- Unsupported audio format or corrupted file
- audio_url not publicly accessible — fetch failure
- num_speakers out of 1-10 range — validation error
- accept_terms_hash missing or invalid — terms not accepted error
- Base64 audio malformed — decoding error
- Network timeout fetching remote audio URL

## How this service works

Transkription mit Sprechertrennung bis 2 Min Audio (pyannote + Whisper), verarbeitet in Deutschland. Speech-to-text with speaker diarization for audio up to 2 min, German-optimised, processed in Germany. Voraussetzung/required: accept_terms_hash, business_name (B2B) – /legal/terms.json

## Output

Returns a transcript broken down by speaker segments, with speaker labels (e.g. SPEAKER_00, SPEAKER_01), timestamps, and the transcribed text for each segment. Processed entirely within Germany for GDPR compliance.

## Request schema (JSON Schema)

```json
{
 "type": "object",
 "properties": {
  "filename": {
   "type": "string",
   "description": "Dateiname mit Endung, z. B. memo.m4a"
  },
  "language": {
   "type": "string",
   "description": "de (Standard), en, auto ..."
  },
  "audio_url": {
   "type": "string",
   "description": "HTTPS-URL der Audiodatei / public URL of the audio file"
  },
  "audio_base64": {
   "type": "string",
   "description": "Audio als Base64 / audio as base64"
  },
  "num_speakers": {
   "type": "integer",
   "description": "Nur /speakers: bekannte Sprecherzahl (1-10), verbessert die Trennung"
  },
  "business_name": {
   "type": "string",
   "description": "Unternehmen, für das gehandelt wird (nur B2B)"
  },
  "privacy_contact": {
   "type": "string",
   "description": "E-Mail für Datenschutz-Meldungen (Art. 33 DSGVO)"
  },
  "accept_terms_hash": {
   "type": "string",
   "description": "SHA-256 der AGB aus /legal/terms.json"
  }
 }
}
```

## More

- Live health (JSON, refreshed every minute): https://www.zero.xyz/c/nexari-speech-to-text-with-speaker-diarization-2-min-ed3d0515/health.json
- [Zero catalog index](https://www.zero.xyz/llms.txt)
- [Other services from x402.nexari.cloud](https://www.zero.xyz/host/x402.nexari.cloud/llms.txt)
