Skip to main content
Speaker diarization answers who spoke when in a recording. Combined with Hear transcription, it produces speaker-labelled text and timestamps for call review, searchable conversations, and downstream analysis. Use Hear async transcription jobs for this workflow. Speak converts text into speech; it does not diarize recordings.

Choose mono diarization or stereo channel separation

Do not set both options. Prefer channel separation when the recording already isolates participants. Two channels containing the same mixed conversation do not provide separate participants. Speaker labels are neutral. They do not identify a real person, establish an agent/customer role, or remain stable across different jobs. Map roles using your own call context before submitting a transcript to Recap.

Submit a mono recording

Use a key with hear:transcribe and transcribe:jobs scopes. Replace the example URL with an HTTPS recording URL that PyAI can fetch.
A 202 response returns a job_id and status: "queued". Poll that job:
Set JOB_ID to the returned ID. Stop polling when status is completed, failed, or cancelled. For a stereo recording with separate participants, replace diarize: true with channel: true.

Check speaker labels before using the result

A completed job does not guarantee that speaker diarization was produced. Read result.segments and check for speaker labels before relying on speaker attribution. For large results, fetch the signed result_url first. If mono diarization cannot be produced, the job can still complete with a transcript and no speaker labels. Your application should handle that fallback explicitly rather than assigning every line to an assumed speaker. Speaker-labelled segments contain text, a speaker label, and start/end offsets in seconds. Available words also carry timing; word-level availability varies by language and formatting. Offsets follow the decoded source-media timeline. See timestamped results for the response shape and exact timing behavior.

Use the transcript in your application

  • Call review: follow speaker turns and jump to timestamps in your recording.
  • Conversation search: return matching text alongside its speaker label and time.
  • Call summaries: establish participant roles, then prepare the transcript for Recap.
  • Spoken output: use Speak when the application needs to synthesize text after processing a conversation.

Limits and data handling

This guide describes async recordings, not live speaker diarization on the Hear stream. Mono speaker attribution is model-derived; review representative audio before relying on it in your workflow. Do not combine trace: true with diarize or channel in the same transcription job; the combination returns 400. See Trace for its separate transcript-scanning workflow. URL-fetched input bytes are not persisted. Uploaded input audio is retained for up to 7 days and job result artifacts for up to 30 days. Read the async jobs guide for upload limits, signed result URLs, idempotency, cancellation, webhooks, and retention details.

Frequently asked questions

Is speaker diarization the same as speaker identification?

No. Diarization groups speaker turns within a recording. It does not verify a speaker’s name or identity across recordings.

Does Speak provide speaker diarization?

Speak generates audio from text. Use Hear async transcription jobs to obtain speaker-labelled transcripts from recordings.

Can I keep separate agent and customer channels?

Yes. Use channel: true when each participant occupies a separate stereo channel. PyAI returns neutral labels; use your recording metadata to map them to the correct roles.

How do I migrate an existing diarization workflow?

Submit recordings to Hear async jobs and adapt polling, fallback handling, speaker labels and timestamps to the PyAI response. Validate representative recordings before moving production traffic. Consult the API reference for the current request and response contract.