Recipes← all recipesView on GitHub

Play a prompt and collect digits or speech over REST

Voiceprompt playback and input collection on a live call

Speak a prompt into a live call from your own backend, stop it early, and collect keypad digits or speech, with the result delivered to your webhook.

restdtmfspeech

The claim

Prompting a caller and reading their answer is usually written into the call’s document. The same two operations exist as REST commands. A process that holds the call id can run them on a call it did not set up. The vendored REST spec, tools/openapi/rest.json, is the authority for the shapes.

Why it holds

  • calling.play requires control_id and play, an ordered list of items. Each item has a type, one of audio, tts, silence or ringtone, and params. A tts item requires text; an audio item requires url.
  • calling.play.stop requires control_id, the id you gave the play.
  • calling.play.volume requires control_id and volume, a dB adjustment the spec bounds to “between -40 and 40”. The recipe refuses anything outside it.
  • calling.collect requires control_id and, in the spec’s words, “at least one of digits or speech”. digits requires max. Results are “delivered asynchronously via the status_url webhook”, so the recipe refuses a collect with no status URL before it is sent.
  • start_input_timers defaults to false, and then initial_timeout does not start counting. Both collects send true. Send false instead when the prompt is still playing, and start the clock yourself with calling.collect.start_input_timers for the same control_id.

The verifier proves the six requests and the refusal. The spec is the authority for what the platform does with them.

How it works

def say(call_id, text, control_id=PLAY_ID, status_url=None):
    item = {"type": "tts", "params": {"text": text}}
    params = {"control_id": control_id, "play": [item]}
    if status_url:
        params["status_url"] = status_url
    return client.calling.play(call_id, **params)

def ask_digits(call_id, status_url, max_digits=10, control_id=COLLECT_ID):
    _needs_status_url(status_url)
    digits = {"max": max_digits, "terminators": "#", "digit_timeout": 5}
    return client.calling.collect(call_id, control_id=control_id, digits=digits,
                                  initial_timeout=10, start_input_timers=True,
                                  status_url=status_url)

What the platform receives for ask_digits:

{"command": "calling.collect",
 "id": "6d3f4a0e-2b1c-4e7a-9f0d-1c2b3a4d5e6f",
 "params": {"control_id": "agent-desk-input",
            "initial_timeout": 10,
            "digits": {"max": 10, "terminators": "#", "digit_timeout": 5},
            "start_input_timers": true,
            "status_url": "https://your-host/collect-events"}}

The SDK’s calling namespace, rest/namespaces/calling.py, sends every call command to that one path and puts the call id in id. Each operation gets its own control id, and the matching stop command names it.

The spec documents what arrives at status_url. A digit result has the shape {control_id, call_id, node_id, result: {type: "digit", params: {digits, terminator}}}; a speech result carries {type: "speech", params: {text, confidence}}. Playback events carry a state of playing, paused, finished or error. The route that receives them is an ordinary webhook handler, the same shape as the one in Handle call status callbacks.

Limitations

The verifier proves the requests, not the audio or the recognition. What the caller hears, and what text the speech engine returns, are live behaviour.

The recipe gives the play and the collect fixed control ids. Two prompts playing at once on one call need two ids; the spec says the id must be unique per active play on the call.

What to change first

Change "tts" in say to "speech" and run the verifier. The item type is outside the spec’s enum and the exact body comparison fails, which is the point: the four item types are the whole vocabulary.