1. Send the called party’s audio
You need a PyAI key withamd:detect and an HTTPS endpoint to receive results.
A sandbox key includes AMD scopes for testing. No PyAI SDK
or separate transcription connection is required for this Twilio integration.
For a direct outbound Twilio call, return this TwiML on the called-party leg:
<Start><Stream> so Twilio continues your call flow while AMD listens.
Always include a following verb; otherwise Twilio ends the call. The test pause
keeps the call open and does not delay the callback. <Connect><Stream> blocks
subsequent TwiML and is intended for two-way audio applications.
On the called-party leg, inbound_track carries audio from that party into
Twilio. In a bridged or conference call, confirm the leg before starting the
stream. Do not send both speakers. Omit language for automatic language detection.
2. Handle the result
Your per-callwebhook receives a JSON POST with decision fields. Example:
call_id is the streamed Twilio
CallSid. Your application controls what happens next:
AMD returns one classification per stream. It does not later send a human
result when someone takes over. A machine result is not an instruction to hang
up, and
voicemail_ready: false does not authorize a voicemail drop.
3. Choose how long to wait
Setdecision_timeout_ms to 3000 or 5000 for a three- or five-second budget.
Allowed range: 1000–15000 ms. The timer starts when PyAI accepts the stream’s
authenticated start, not when Twilio answers the call. Results can arrive earlier;
a shorter budget may produce more unknowns. Allow additional time for webhook delivery.
aggressiveness is optional. Start with 0.25 when protecting human pickups is
the priority. Omit language for multilingual calls. Pass correlation values,
such as lead_id, as string parameters; they return under custom_parameters.
Reference
The sections below cover direct WebSocket clients, optional fields and account configuration. Twilio users can start with the three steps above.Decision cutoff and default timing
Decision cutoff and default timing
Set
decision_timeout_ms on each stream to 3000 for three seconds or 5000
for five seconds. Twilio sends it as a <Parameter> as shown above. Clients
that send Media Streams frames directly put it in start.customParameters:- Accepts an integer or decimal integer string from 1000 to 15000 milliseconds, in every supported language. It is a per-stream option, not an account config field.
- The clock begins when PyAI accepts the authenticated start frame and valid parameters, before recognizer startup. It is not measured from carrier answer.
- A decisive result returns earlier. At the cutoff, AMD uses evidence already
received by that time and otherwise returns
unknownwithrule_id: "decision_timeout". It does not force a human or machine guess. - Stalled audio and a pending final transcript share the same budget. The cutoff does not advance audio time or manufacture silence. A shorter budget may produce more unknown results.
decision_elapsed_msreports elapsed time to result preparation;decision_timeout_msechoes the requested budget. Webhook transit and receiver acknowledgement take additional time, so allow for delivery in your fallback timer.
decision_timeout_ms to preserve the default behavior. The separate
decision_window_ms option limits processed audio, accepts 1000–15000 ms,
and is available for English streams. Either limit can finish a decision first.
The default English audio window is five seconds and can extend by up to two
seconds for eligible calls; that extension never extends an explicit elapsed cutoff.Invalid cutoff values produce an error event with code
invalid_stream_parameters and close the socket with code 1008.Result fields and stored records
Result fields and stored records
A
machine decision can be further classified by subtype, voicemail, ivr,
screening (iPhone/Google Call Screen), or music (hold music). The wire event includes subtype; the stored record folds that subtype into
answered_by. Read the record with
GET /v1/amd/calls/{id} (or on the amd.call.completed webhook below), where
answered_by is one of human, machine, voicemail, live_voicemail,
ivr, screening, music, human_gatekeeper, sit_invalid, fax, silence,
or unknown. A generic machine result preserves uncertainty about the kind of
automated answer. It does not mean a voicemail greeting has finished, and it does
not trigger managed-call automatic hangup.Timing diagnostics appear at the top level of socket events, both webhook
payloads, and stored call details. Stored details also include them under
meta.decision. They exclude webhook transport and receiver acknowledgement.The WebSocket message adds event: "amd"; the per-call HTTP callback has the decision fields without that envelope.Webhooks and custom parameters
Webhooks and custom parameters
Two webhook paths, carrying different payloads:Use the per-call callback for routing. The account event describes AMD completion, not the end of the Twilio phone call. If both callbacks are enabled, deduplicate before taking a routing action.This example shows an idle stream: zero audio processed despite three seconds
elapsed. The payload excerpt omits other result fields. Correlation fields stay
nested and cannot replace the AMD verdict or call ID.Names must match Reject stale timestamps and deduplicate retries by call/event id. Mint or rotate
the organization secret with
- Per-call, the TwiML
<Parameter name="webhook">. The moment the decision lands, PyAI POSTs the decision fields (the coarseanswered_byclass, without the socket’sevent: "amd"envelope) to that URL, for that call only. - Account-wide,
webhook_urlinPOST /v1/amd/config. When the AMD record completes, PyAI POSTs a signedamd.call.completedevent with the full call record, including the machine subtype inanswered_by:
Return your own correlation parameters
Pass string fields such aslead_id and campaign_id alongside the AMD
parameters in start.customParameters / TwiML <Parameter> entries. They return
under custom_parameters in the socket event, per-call webhook, account completion
webhook and stored call detail:[A-Za-z0-9][A-Za-z0-9_.-]{0,63}. Send at most 32 fields,
1024 UTF-8 bytes per string value, and 8192 combined UTF-8 bytes for all names
and values. Invalid fields cause invalid_stream_parameters. Credentials,
internal names beginning with _, and reserved configuration parameters are
excluded. Do not send secrets as correlation fields. Query parameters in the
webhook URL remain on that URL; they are not copied into the JSON body.For the account-wide completion webhook, verify
X-PyAI-Signature: t=<unix_seconds>,v1=<hex> against the exact raw body:POST /v1/webhooks/signing-secret. The per-call
TwiML webhook is a separate low-latency callback; do not assume it has the same
full-record payload as amd.call.completed.The per-call callback sends JSON and does not carry the account webhook signature
or X-Twilio-Signature. Protect that receiver with your own validation; for an
authenticated result, use the signed account event or fetch the stored call
with your PyAI key. A callback response acknowledges delivery; returning TwiML
to PyAI does not change the Twilio call.Audio and language
Audio and language
Send only the called party’s audio, paced in real time. Do not mix the
agent’s speech into the stream. When forking a Twilio call, attach the stream
to the called-party leg and verify which track contains that party; a track
name alone does not identify the person being called.AMD combines audio evidence with accumulated recognition from Hear streaming.
Omit the stream’s
language parameter for automatic language detection. An
explicit supported language hint can constrain recognition; do not hard-code
en for multilingual calls. A separate customer STT connection is not required
for AMD. You can still open Hear separately if your application needs transcripts.A recognizable screening prompt can produce an early result, but a short
“hello,” background noise, or a pause alone does not prove who answered.
Recent automation evidence can defer a pause-based human decision. That does
not force a machine verdict: conflicting or insufficient evidence may remain
unknown. Existing clients receive these server-side improvements automatically.Twilio Stream URLs do not support query parameters; use nested <Parameter> entries. For server-side WebSocket clients, use the pyai.v1 and pyai-key.<API_KEY> subprotocol pair.Aggressiveness
Aggressiveness
AMD has a single operating-point dial,
aggressiveness ∈ [0, 1], set per account
(POST /v1/amd/config) or per call (a TwiML <Parameter>):- 0.0-0.25, human-safe (default). Prioritizes avoiding false machine decisions. For predictive
dialers with live agents; on the deadline it returns
unknown(let the agent listen) rather than risk a falsemachine. - 0.6-1.0, more aggressive timing. Adjusts evidence timing thresholds. It does not force a machine guess or establish voicemail recording readiness.
unknown at every setting. Confidence is a rule
score, not a calibrated probability or a guarantee against classification errors.The error costs are asymmetric, a false machine (hang up on a prospect) is far
worse than a false human (waste a few agent-seconds), so you pick your point on
the curve instead of living with one fixed default.Account configuration and SDK examples
Account configuration and SDK examples
API and billing
Stream:wss://api.pyai.com/v1/amd/stream (amd:detect). Configure account
defaults: POST /v1/amd/config (amd:configure). Read results:
GET /v1/amd/calls and GET /v1/amd/calls/{id} (amd:read).
AMD records one amd.calls unit per answered call. See pricing
for rates and included usage.