The supervisor inbox is a single queue of live calls that a human supervisor may want to take over. Two things put a call there: an escalation raised by the agent runtime, and a supervisor manually picking a call to join. The screen body is one shared implementation (SupervisorInboxBody, from @telequick/agent-ui/screens/supervisor-inbox). The contact-centre console and the voice-AI console both render it, so the queue behaves identically in either place. Only the chrome differs — each console supplies its own admin rail, page header, and help link:
If you are reading this from one console and comparing notes with someone in the other, you are looking at the same component.

What puts a call in the inbox

Both kinds of item are only actionable while the call is up. Once the call ends, what remains is the Call Detail Record, not an inbox item — see Telemetry.

Escalation signals from the agent runtime

Escalations come from the agent DAG, not from this screen. The screen is a consumer: it renders whatever the runtime has flagged. That means the conditions under which a call gets flagged are a property of the agent you deployed, and changing them is an agent-configuration change, not a console setting. There is no platform-wide escalation threshold to look up here. Calls that routed through an agent DAG are identifiable after the fact: the CDR row for such a call has agent_id set (empty for calls that never touched an agent). Pair that with call_sid to correlate an inbox item with its recorded call.

Join modes: listen, whisper, and barge

Three modes describe how much of the supervisor’s audio reaches the call. These are the standard industry meanings, and they are what the mode selector on this screen is choosing between: The mode is chosen at join time. Escalated items and manually selected items use the same three modes; a flag from the agent runtime does not force a particular one.

Audio routing and the mixer dependency

Audio routing follows the chosen mode once the codec_engine mixer is live. The mixer is what makes a third party audible on an existing two-party call, so a join mode has no audio effect until the mixer is serving that call’s shard. Practical consequences while the mixer is not live for a deployment:
  • The inbox, the flags, and the mode selection all still render and work as a control-plane surface.
  • Do not treat a successful join as proof that a supervisor can hear or be heard. Confirm on the audio path itself.
The mixer exposes its own health signal. Watch the telequick_codec_engine_lag_us histogram (labelled by shard) alongside the per-tenant audio-quality gauges (telequick_rtp_loss_ratio, telequick_jitter_ms, telequick_estimated_mos) when investigating a join that produced no audio. For definitions of those quality numbers, see Telephony Metrics.

Verbs and events behind the join

The join action is not a new control surface. It sits on the existing voice control plane:
  • barge is the call-control verb on the root client (TeleQuickClient), alongside dial, hangup, and push_audio. It is the verb that attaches a party to a call already in progress.
  • push_audio pushes audio into an existing call, addressed by call_sid.
  • CallEvent is the event stream you observe the call on. The terminal event is CHANNEL_HANGUP_COMPLETE, which carries the same fields the gateway persists to the CDR. When that event arrives for a call, its inbox item is no longer joinable.
Every one of these RPCs produces an OpenTelemetry span carrying tenant.id, call.sid, method.id, and transport, so a join is traceable end to end from the console action to the SIP/RTP setup.

Permissions and audit

A supervisor join crosses both auth surfaces described in Authentication:
  • The console’s control-plane call authenticates with an API key scoped to an org and a set of capabilities. The control-plane API rejects a key without the capability scope for the procedure it is calling, so a console that cannot issue call control cannot join.
  • The supervisor’s own audio session is data plane. It needs a short-lived relay token whose ns claim covers the call’s namespace (voice/<callSid>/*). The relay’s namespace_auth hook checks that token on every publish and subscribe, which is what scopes a supervisor to one call rather than to the tenant’s whole voice traffic. Mint the token server-side; never ship the API key to the console.
For reconstructing who joined what, the durable records are the telemetry streams rather than anything this screen stores: the telequick_admin_requests_total counter (labelled method, result) covers control-plane admin calls, traces carry call.sid and tenant.id, and the call’s CDR row carries call_sid, tenant_id, agent_id, and recording_url.