Choose mono diarization or stereo channel separation
Do not set both options. Prefer channel separation when the recording already
isolates participants. Two channels containing the same mixed conversation do
not provide separate participants.
Speaker labels are neutral. They do not identify a real person, establish an
agent/customer role, or remain stable across different jobs. Map roles using
your own call context before submitting a transcript to Recap.
Submit a mono recording
Use a key withhear:transcribe and transcribe:jobs scopes. Replace the example
URL with an HTTPS recording URL that PyAI can fetch.
202 response returns a job_id and status: "queued". Poll that job:
JOB_ID to the returned ID. Stop polling when status is completed,
failed, or cancelled. For a stereo recording with separate participants,
replace diarize: true with channel: true.
Check speaker labels before using the result
A completed job does not guarantee that speaker diarization was produced. Readresult.segments and check for speaker labels before relying on speaker
attribution. For large results, fetch the signed result_url first.
If mono diarization cannot be produced, the job can still complete with a
transcript and no speaker labels. Your application should handle that fallback
explicitly rather than assigning every line to an assumed speaker.
Speaker-labelled segments contain text, a speaker label, and start/end
offsets in seconds. Available words also carry timing; word-level availability
varies by language and formatting. Offsets follow the decoded source-media
timeline. See timestamped results
for the response shape and exact timing behavior.
Use the transcript in your application
- Call review: follow speaker turns and jump to timestamps in your recording.
- Conversation search: return matching text alongside its speaker label and time.
- Call summaries: establish participant roles, then prepare the transcript for Recap.
- Spoken output: use Speak when the application needs to synthesize text after processing a conversation.
Limits and data handling
This guide describes async recordings, not live speaker diarization on the Hear stream. Mono speaker attribution is model-derived; review representative audio before relying on it in your workflow. Do not combinetrace: true with diarize or channel in the same transcription
job; the combination returns 400. See Trace for its
separate transcript-scanning workflow.
URL-fetched input bytes are not persisted. Uploaded input audio is retained for
up to 7 days and job result artifacts for up to 30 days. Read the
async jobs guide for upload limits, signed
result URLs, idempotency, cancellation, webhooks, and retention details.
Frequently asked questions
Is speaker diarization the same as speaker identification?
No. Diarization groups speaker turns within a recording. It does not verify a speaker’s name or identity across recordings.Does Speak provide speaker diarization?
Speak generates audio from text. Use Hear async transcription jobs to obtain speaker-labelled transcripts from recordings.Can I keep separate agent and customer channels?
Yes. Usechannel: true when each participant occupies a separate stereo
channel. PyAI returns neutral labels; use your recording metadata to map them
to the correct roles.