Skip to main content
PyAI AMD listens to the person you are calling and reports human, machine, screening, invalid number or unknown. It returns as soon as it has decisive evidence, or returns uncertainty at your chosen cutoff. Already using Twilio AMD? Follow Replace Twilio AMD for the before-and-after setup and callback changes. You keep Twilio for calling.

1. Send the called party’s audio

You need a PyAI key with amd:detect and an HTTPS endpoint to receive results. A sandbox key includes AMD scopes for testing. No PyAI SDK or separate transcription connection is required for this Twilio integration. For a direct outbound Twilio call, return this TwiML on the called-party leg:
Use <Start><Stream> so Twilio continues your call flow while AMD listens. Always include a following verb; otherwise Twilio ends the call. The test pause keeps the call open and does not delay the callback. <Connect><Stream> blocks subsequent TwiML and is intended for two-way audio applications. On the called-party leg, inbound_track carries audio from that party into Twilio. In a bridged or conference call, confirm the leg before starting the stream. Do not send both speakers. Omit language for automatic language detection.

2. Handle the result

Your per-call webhook receives a JSON POST with decision fields. Example:
This is an excerpt, not the complete payload. call_id is the streamed Twilio CallSid. Your application controls what happens next: AMD returns one classification per stream. It does not later send a human result when someone takes over. A machine result is not an instruction to hang up, and voicemail_ready: false does not authorize a voicemail drop.

3. Choose how long to wait

Set decision_timeout_ms to 3000 or 5000 for a three- or five-second budget. Allowed range: 1000–15000 ms. The timer starts when PyAI accepts the stream’s authenticated start, not when Twilio answers the call. Results can arrive earlier; a shorter budget may produce more unknowns. Allow additional time for webhook delivery. aggressiveness is optional. Start with 0.25 when protecting human pickups is the priority. Omit language for multilingual calls. Pass correlation values, such as lead_id, as string parameters; they return under custom_parameters.

Reference

The sections below cover direct WebSocket clients, optional fields and account configuration. Twilio users can start with the three steps above.
Set decision_timeout_ms on each stream to 3000 for three seconds or 5000 for five seconds. Twilio sends it as a <Parameter> as shown above. Clients that send Media Streams frames directly put it in start.customParameters:
  • Accepts an integer or decimal integer string from 1000 to 15000 milliseconds, in every supported language. It is a per-stream option, not an account config field.
  • The clock begins when PyAI accepts the authenticated start frame and valid parameters, before recognizer startup. It is not measured from carrier answer.
  • A decisive result returns earlier. At the cutoff, AMD uses evidence already received by that time and otherwise returns unknown with rule_id: "decision_timeout". It does not force a human or machine guess.
  • Stalled audio and a pending final transcript share the same budget. The cutoff does not advance audio time or manufacture silence. A shorter budget may produce more unknown results.
  • decision_elapsed_ms reports elapsed time to result preparation; decision_timeout_ms echoes the requested budget. Webhook transit and receiver acknowledgement take additional time, so allow for delivery in your fallback timer.
Omit decision_timeout_ms to preserve the default behavior. The separate decision_window_ms option limits processed audio, accepts 1000–15000 ms, and is available for English streams. Either limit can finish a decision first. The default English audio window is five seconds and can extend by up to two seconds for eligible calls; that extension never extends an explicit elapsed cutoff.Invalid cutoff values produce an error event with code invalid_stream_parameters and close the socket with code 1008.
A machine decision can be further classified by subtype, voicemail, ivr, screening (iPhone/Google Call Screen), or music (hold music). The wire event includes subtype; the stored record folds that subtype into answered_by. Read the record with GET /v1/amd/calls/{id} (or on the amd.call.completed webhook below), where answered_by is one of human, machine, voicemail, live_voicemail, ivr, screening, music, human_gatekeeper, sit_invalid, fax, silence, or unknown. A generic machine result preserves uncertainty about the kind of automated answer. It does not mean a voicemail greeting has finished, and it does not trigger managed-call automatic hangup.Timing diagnostics appear at the top level of socket events, both webhook payloads, and stored call details. Stored details also include them under meta.decision. They exclude webhook transport and receiver acknowledgement.The WebSocket message adds event: "amd"; the per-call HTTP callback has the decision fields without that envelope.
Two webhook paths, carrying different payloads:
  • Per-call, the TwiML <Parameter name="webhook">. The moment the decision lands, PyAI POSTs the decision fields (the coarse answered_by class, without the socket’s event: "amd" envelope) to that URL, for that call only.
  • Account-wide, webhook_url in POST /v1/amd/config. When the AMD record completes, PyAI POSTs a signed amd.call.completed event with the full call record, including the machine subtype in answered_by:
Use the per-call callback for routing. The account event describes AMD completion, not the end of the Twilio phone call. If both callbacks are enabled, deduplicate before taking a routing action.

Return your own correlation parameters

Pass string fields such as lead_id and campaign_id alongside the AMD parameters in start.customParameters / TwiML <Parameter> entries. They return under custom_parameters in the socket event, per-call webhook, account completion webhook and stored call detail:
This example shows an idle stream: zero audio processed despite three seconds elapsed. The payload excerpt omits other result fields. Correlation fields stay nested and cannot replace the AMD verdict or call ID.Names must match [A-Za-z0-9][A-Za-z0-9_.-]{0,63}. Send at most 32 fields, 1024 UTF-8 bytes per string value, and 8192 combined UTF-8 bytes for all names and values. Invalid fields cause invalid_stream_parameters. Credentials, internal names beginning with _, and reserved configuration parameters are excluded. Do not send secrets as correlation fields. Query parameters in the webhook URL remain on that URL; they are not copied into the JSON body.For the account-wide completion webhook, verify X-PyAI-Signature: t=<unix_seconds>,v1=<hex> against the exact raw body:
Reject stale timestamps and deduplicate retries by call/event id. Mint or rotate the organization secret with POST /v1/webhooks/signing-secret. The per-call TwiML webhook is a separate low-latency callback; do not assume it has the same full-record payload as amd.call.completed.The per-call callback sends JSON and does not carry the account webhook signature or X-Twilio-Signature. Protect that receiver with your own validation; for an authenticated result, use the signed account event or fetch the stored call with your PyAI key. A callback response acknowledges delivery; returning TwiML to PyAI does not change the Twilio call.
Send only the called party’s audio, paced in real time. Do not mix the agent’s speech into the stream. When forking a Twilio call, attach the stream to the called-party leg and verify which track contains that party; a track name alone does not identify the person being called.AMD combines audio evidence with accumulated recognition from Hear streaming. Omit the stream’s language parameter for automatic language detection. An explicit supported language hint can constrain recognition; do not hard-code en for multilingual calls. A separate customer STT connection is not required for AMD. You can still open Hear separately if your application needs transcripts.A recognizable screening prompt can produce an early result, but a short “hello,” background noise, or a pause alone does not prove who answered. Recent automation evidence can defer a pause-based human decision. That does not force a machine verdict: conflicting or insufficient evidence may remain unknown. Existing clients receive these server-side improvements automatically.Twilio Stream URLs do not support query parameters; use nested <Parameter> entries. For server-side WebSocket clients, use the pyai.v1 and pyai-key.<API_KEY> subprotocol pair.
AMD has a single operating-point dial, aggressiveness ∈ [0, 1], set per account (POST /v1/amd/config) or per call (a TwiML <Parameter>):
  • 0.0-0.25, human-safe (default). Prioritizes avoiding false machine decisions. For predictive dialers with live agents; on the deadline it returns unknown (let the agent listen) rather than risk a false machine.
  • 0.6-1.0, more aggressive timing. Adjusts evidence timing thresholds. It does not force a machine guess or establish voicemail recording readiness.
Uncertain evidence can return unknown at every setting. Confidence is a rule score, not a calibrated probability or a guarantee against classification errors.The error costs are asymmetric, a false machine (hang up on a prospect) is far worse than a false human (waste a few agent-seconds), so you pick your point on the curve instead of living with one fixed default.

API and billing

Stream: wss://api.pyai.com/v1/amd/stream (amd:detect). Configure account defaults: POST /v1/amd/config (amd:configure). Read results: GET /v1/amd/calls and GET /v1/amd/calls/{id} (amd:read). AMD records one amd.calls unit per answered call. See pricing for rates and included usage.