Recipes← all recipesView on GitHub

Stream call audio to a WebSocket that checks a bearer token

Voiceauthenticated media streaming over WebSocket

Send one side of a call, or both, to your own WebSocket over TLS, with a bearer token on the connection and your own metadata attached.

restmedia-streaming

The claim

tap copies a call’s audio to a URI. calling.stream does the same job for an endpoint that has opinions. It insists on TLS and authenticates. It lets you pick a side of the conversation, and it hands your endpoint metadata when the connection opens.

Why it holds

The vendored REST spec, tools/openapi/rest.json, is the authority.

  • calling.stream requires control_id and url. The url “must start with wss:// (TLS is required; plain ws:// is rejected)“, so the recipe refuses a ws:// url before sending one.
  • track is inbound_track, outbound_track or both_tracks.
  • authorization_bearer_token is “included as Authorization: Bearer <token> when establishing the WebSocket connection”.
  • custom_parameters is an “arbitrary JSON object passed through to the WebSocket endpoint as connection metadata”, which is how a stream says which ticket it belongs to.
  • status_url receives the lifecycle webhooks, with status_url_method of GET or POST.
  • calling.stream.stop requires control_id.

Use tap when the destination is an RTP address or a plain socket you control. Use stream when the endpoint authenticates the connection, or when only one side of the call should leave the platform.

How it works

def start(call_id, url, track="both_tracks", control_id=CONTROL_ID,
          status_url=None, tag=None):
    if not url.startswith("wss://"):
        raise ValueError(...)
    params = {"control_id": control_id, "url": url, "track": track,
              "codec": "PCMU", "name": "support"}
    if STREAM_TOKEN:
        params["authorization_bearer_token"] = STREAM_TOKEN
    if tag:
        params["custom_parameters"] = {"tag": tag}
    return client.calling.stream(call_id, **params)

What the platform receives:

{"command": "calling.stream",
 "id": "6d3f4a0e-2b1c-4e7a-9f0d-1c2b3a4d5e6f",
 "params": {"control_id": "support-audio",
            "url": "wss://media.example.com/calls",
            "track": "both_tracks",
            "codec": "PCMU",
            "name": "support",
            "authorization_bearer_token": "…",
            "custom_parameters": {"tag": "ticket-4417"},
            "status_url": "https://media.example.com/stream-events",
            "status_url_method": "POST"}}

The token is read from the environment, so it is never a literal in the code the page shows. Your endpoint checks the Authorization header on the upgrade request and refuses anything else, which is the point of sending it.

codec is freeform in the spec, which lists PCMU, PCMA and OPUS as common values. What your endpoint accepts is the real constraint.

Limitations

The verifier proves the requests, not the socket. Whether your endpoint accepts the upgrade, and what it does with the audio, is yours to build.

The bearer token is sent to the platform in the clear inside the command body, which is a TLS request to the API. Treat it as a credential your endpoint issues, and rotate it like one.

What to change first

Point start at a ws:// url and run the verifier. It raises before the request, which is the point. The spec rejects plain WebSocket, and the check belongs in your code rather than in a 400.