# Agent orchestrator and live agent events

> How the orchestrator screen binds trunks to agents, what each live agent event line means, and how per-trunk session counts are derived.

The orchestrator screen is the operator view of which trunk hands its
calls to which agent, plus a rolling tail of what the agent runtime is
emitting right now. It combines three sources: the `trunk` table for the
static binding, `telephony.listActiveCalls` for live per-trunk session
counts, and `monitoring.agentEvents` for the event tail.

Agent orchestration itself runs in the native C++/Seastar `agent_runtime`
binary. The console does not drive dispatch; it reports on it.

## What the orchestrator screen shows

The screen has three regions, top to bottom:

| Region | Source | Refresh |
| ------ | ------ | ------- |
| **Live Agent Events** | `monitoring.agentEvents` | Polled every 2 s, pausable |
| **Capacity KPIs** | `trunk` rows for the active org | On load / manual refresh |
| **Trunk Orchestration** table | `trunk` rows + `telephony.listActiveCalls` | Bindings on refresh; session counts every 10 s |

The four KPI tiles are derived purely from trunk configuration, not from
live traffic:

- **Configured Trunks** — the number of `trunk` rows for the active org.
- **Total Channel Cap** — the sum of `channel_limit` across those rows.
- **Trunks w/ Agent** — the count of rows with a non-null `agent_id`.
- **Avg Channels/Trunk** — total channel cap divided by trunk count,
  rounded.

None of these tiles move when calls start or stop. They change when you
change trunk configuration.

## How a trunk is bound to an agent

A binding is a single column. Each `trunk` row carries an `agent_id`; if
it is set, inbound calls arriving on that trunk are dispatched to that
agent by the runtime. If it is null, the table renders `none` in the
Agent column and the trunk contributes nothing to the **Trunks w/ Agent**
tile.

The table renders these fields per trunk:

| Column | Field | Notes |
| ------ | ----- | ----- |
| Trunk | `display_name`, `trunk_id` | The name falls back to `—` when unset; the `trunk_id` renders below it as monospace. |
| Direction | `direction` | `1` renders as **Outbound**; anything else renders as **Inbound**. |
| Channel Limit | `channel_limit` | The configured cap, verbatim. |
| Agent | `agent_id` | The bound agent, or `none`. |
| Sessions | derived | See [How per-trunk session counts are derived](#how-per-trunk-session-counts-are-derived). |

Because the binding is per trunk rather than per call, an agent can be
bound to several trunks at once, and a trunk is bound to at most one
agent.

If the org has no trunks at all, the table is replaced with an empty
state prompting you to add a trunk and bind it to an agent.

## Reading a live agent event line

The **Live Agent Events** panel polls `monitoring.agentEvents` for the
per-tenant ring of telemetry records the agent runtime publishes on its
`telequick.agent.*` topics. Polling is used deliberately rather
than a push transport: the control-plane API stays stateless from the
browser's point of view, and closing the tab simply stops the next poll.

Each row is laid out in fixed columns:

```
14:22:07  rtp   sid=sess_a1b2c3d4   { … remaining JSON fields … }
```

| Column | Field | Meaning |
| ------ | ----- | ------- |
| Time | `receivedAtMs` | When the control-plane API received the record, rendered as local `HH:MM:SS`. This is receive time, not the runtime's emit time. |
| Channel | `channel` | Which `telequick.agent.*` topic the record arrived on. This is the short channel suffix, so you can tell at a glance which part of the runtime spoke. |
| Session | `sessionId` | Rendered as `sid=…` and truncated to fit. Omitted entirely when the record carries no session. Use it to correlate lines belonging to one call. |
| Payload | `raw` | The record body, pretty-printed inline. |

The payload column hides `session_id` and `tenant_id`, because the
session is already shown in its own column and the tenant is implied by
the org you are viewing. Everything else in the record is rendered as-is.

Newest events render at the top. The panel buffers up to the query's
`limit` (the screen requests 100) and reports the buffered count in the
header. **Pause** stops the 2 s refresh so a line does not scroll away
while you read it; **Resume** restarts it. Polling also stops while the
tab is in the background.

`monitoring.agentEvents` returns an `enabled` flag alongside the events.
When the telemetry consumer is not enabled, the procedure returns an
empty event list. An empty panel with a live broker shows a prompt to
run **Test Drive** on an agent to produce a first event.

## How per-trunk session counts are derived

The **Sessions** column is not reported by the runtime. The screen calls
`telephony.listActiveCalls` for the org every 10 seconds — the same
cadence the dashboard's active-calls view uses — and groups the returned
rows by their `trunk_id` client-side. A trunk with at least one matching
row renders a pulsing dot and `N live`, suffixed with `/ channel_limit`
when a channel limit is configured. A trunk with no matching rows renders
a muted dot and `0 live`.

`telephony.listActiveCalls` builds its result like this:

1. Read call SIDs from the per-org active-calls sorted set,
   `telequick:<orgId>:active_calls`, highest score first, bounded
   by the `limit` input.
2. For each candidate SID, check whether the runtime has written a
   `session:<sid>:end` hash. The runtime writes that key on hangup, via
   the RTP listener's silence-finalize path or dialog closure. Its
   presence means the call is over.
3. SIDs whose `:end` key exists are treated as stale, removed from the
   listing, and evicted from the sorted set inline. Eviction is
   best-effort; if it fails, the same entries are filtered again on the
   next listing.
4. For the surviving SIDs, read `tenant_id`, `agent_id`, `trunk_id`,
   `realm`, `started_at_ms`, and `node_id` from the `session:<sid>` hash.
5. Drop any row whose `tenant_id` does not match the requested org, and
   any row whose `started_at_ms` is missing, unparseable, or older than
   the window cutoff.
6. Sort by `started_at_ms` descending and truncate to `limit`.

Rows with an empty `trunk_id` are skipped when grouping, so they
contribute to no trunk's count. `realm` defaults to `internal` and
`agent_id`, `trunk_id`, and `node_id` default to empty strings when the
hash field is absent.

Two consequences are worth internalising. First, the count is of sessions
the runtime still considers live, which can differ from what a carrier
believes is up — a carrier-side disconnect that never reached the runtime
leaves a session listed until its `:end` key appears or it ages out of
the window. Second, the `limit` is a ceiling on the listing, so on a very
busy org the per-trunk counts reflect only the most recently started
calls that fit under it.

## Stale session entries and the 60-minute active window

`telephony.listActiveCalls` treats "active" as "started within the last
60 minutes". Anything older is excluded from the sorted-set read and
excluded again by the `started_at_ms` check after the hash fetch.

The window exists because real calls finish in seconds to minutes. An
entry older than the window is almost certainly a sorted-set row that a
hangup path failed to remove — a carrier disconnect, a gateway restart
mid-call, or a Redis blip. The previous 8-hour window was a holdover from
an earlier SCAN-based implementation and surfaced hours-old hung-up calls
in the operator UI.

Stale entries are therefore handled twice over: the `:end` liveness gate
catches calls that ended cleanly but were never removed from the index,
and the window catches everything else. Both paths self-heal — the
`:end` gate evicts what it finds, and the window simply stops reading
old entries.

If a trunk shows `0 live` while you believe a call is up, check whether
the call started more than 60 minutes ago, and whether a
`session:<sid>:end` hash exists for it.

## Metrics not yet exposed by the runtime

The `agent_runtime` binary does not currently publish live runtime
gauges. Specifically, these are **absent**, not zero:

- Active session counts as reported by the runtime itself. The
  orchestrator derives session counts from the active-calls index
  instead, as described above.
- Reactor stalls.
- Per-trunk throughput.

Do not read a blank or missing runtime counter as a healthy zero. Until
the runtime's telemetry sink writes gauges out, the orchestrator shows
the static trunk-to-agent binding plus the derived session counts, and
nothing else about runtime internals.

For the metric families the gateway *does* export today, see
[Telemetry](/platform/telemetry).

## Related

- [Telemetry](/platform/telemetry) — gateway metrics, traces, and CDRs
- [Telephony Metrics](/glossary/metrics) — definitions for channel caps, concurrency, and trunk utilisation
