/app/observability)
is a layer-4 view of the QUIC transport, scoped to one tenant. It answers
four questions: how many connections are open, how much data is moving, what
the round-trip latency looks like, and how often connections survived a
network change.
It deliberately shows nothing else. Voice latency, robotics frame rates, and
games ping are per-modality concerns and live on each vertical’s own portal
(the agent, robotics, and games subdomains). This page explains what the
numbers on the transport screen mean and how they are computed, because the
tiles and the charts do not all use the same aggregation.
What the screen covers
All data comes from a single procedure,observability.quicTransport, which
takes the org id and a window length in minutes and returns two things:
kpi— the six numbers rendered as tiles across the top.series— one row per 1-minute bucket, rendered as the four charts.
tenant_id filter on every query, so a reader
only ever sees their own transport traffic.
What a connection (subject) is
The Connections number iscountDistinct(subject_id) within a bucket, not
a count of sessions opened or an event count.
A subject is the transport-level identity of one QUIC connection as the
relay records it. Every client that terminates a QUIC connection against a
relay contributes one subject, regardless of which modality it belongs to. In
practice the count mixes:
- voice calls,
- robot links,
- spectator viewers.
How the 1-minute buckets are built
Two ClickHouse rollups back the screen. Both are pre-bucketed at one-minute granularity and both are filtered tobucket > now() - INTERVAL <window> MINUTE:
The two result sets are then joined in the API by bucket timestamp: the RTT
rows are indexed by bucket and folded into the connection/throughput rows.
Consequences worth knowing:
- The bucket list comes from
quic_cwnd_1m_q. If a bucket exists in the RTT rollup but has no connection/byte row, it does not appear on the screen. - If a bucket has connection rows but no matching RTT row, the p50/p95/p99
for that bucket are reported as
0. A dip to zero on the latency chart means “no RTT sample for that minute”, not “instant delivery”. - Buckets are emitted only when there is traffic, so gaps in the charts are gaps in data rather than interpolated values.
0 KB/s on the throughput chart while the bytes series
in the fourth chart is clearly non-zero.
RTT percentiles and how they are averaged
The percentiles are not recomputed from raw samples at query time. They are already stored per connection-level rollup row asrtt_p50, rtt_p95, and
rtt_p99; the query takes the arithmetic mean of those percentiles across all
rows in the bucket and rounds to the nearest millisecond.
That means the charted “p99” is an average of per-connection p99s, not the 99th
percentile of the tenant’s whole population of RTT samples. It is the right
shape for spotting a trend or a regression across the fleet, and it is
deliberately insensitive to a single pathological connection: one bad link
drags the averaged p99 up only in proportion to how many connections are in the
bucket.
The subtitle on the chart reads “p50 / p95 / p99 across all connections (ms)”
for this reason — all three lines are fleet-wide roll-ups.
The two RTT tiles are colour-coded so you can triage at a glance. The console
uses these cut-offs for the tile colour only; they are display bands, not a
service commitment:
Path migration: what counts as one
QUIC identifies a connection by its connection ID rather than by the four-tuple, so a client that changes network path keeps the same connection instead of reconnecting. Each time that happens, the transport records a migration, and the rollup’smigrations column carries the count for the
bucket.
The screen reports migrations two ways:
- the Path migrations tile sums
migrationsacross every bucket in the selected window; - the fourth chart plots per-bucket
migrationson the left axis, with rawbytes_in/bytes_outon the right axis as a reference trace.
Reading the KPI tiles vs the charts
The six tiles do not all use the same aggregation, and the hint text under each tile tells you which one it is:
Throughput is averaged on purpose. It is computed as total bytes over the
window divided by
series.length × 60 seconds, so a 30-minute view shows the
steady-state rate rather than whatever the most recent minute happened to do.
Bursty workloads would otherwise make the tile unreadable.
Connections and RTT come from the latest bucket because they are
point-in-time facts: you want to know how many connections are up now and
what latency they see now, not what the average was half an hour ago.
Two practical consequences:
- Changing the window changes the Ingress/Egress and Path-migration tiles, but leaves Connections and the RTT tiles alone (the latest bucket is the same regardless of how far back you look).
- The Ingress/Egress tiles will not match the right-hand end of the throughput chart. The tile is the window average; the chart point is that single bucket’s rate. If they diverge sharply, the traffic is bursty.
p95 value is present in the series and drawn on the
latency chart, but there is no p95 tile — read it off the chart.
Time window and refresh
The window selector offers 15 minutes, 30 minutes, 1 hour, 6 hours, and 24 hours. The default is 30 minutes. The procedure accepts any integer number of minutes from 1 to 1440, so the selector’s options are a subset of what the API allows. The chosen window is mirrored into the?w=<minutes> query parameter via
history.replaceState, which means:
- reloading the page keeps your window,
- and you can deep-link a specific window to a colleague.
w falls back to 30 minutes.
The query refetches every 30 seconds while the screen is open, and the refresh
button in the header forces an immediate refetch. Since the rollups are
1-minute buckets, a manual refresh mid-minute will usually return the same
latest bucket.
Longer windows do not change the bucket size — a 24-hour view is 1-minute
buckets across 24 hours, so expect a dense series.
Empty and error states
The procedure returns the same zeroed shape in two different situations:- No traffic in the window — the rollups return no rows for the tenant, so
seriesis empty and everykpifield is0(withwindow_minutesechoed back). - The ClickHouse query failed — the handler catches the error and returns the same empty payload rather than surfacing an exception.
Where per-modality metrics live instead
This screen stops at the transport layer. Anything that requires knowing what the bytes were belongs to a modality:- Voice — call latency, answer rates, audio quality: the agent portal, and Telephony Metrics.
- Robotics — frame rates and control-loop timing: the robotics portal.
- Games — ping and room health: the games portal.
mod_relay’s counters as Prometheus
text; the client derives rates by differencing successive samples.
That call is carried over the engine’s RPC service (relay.metrics), not over
HTTP. mod_relay does expose a /metrics route, but it is unreachable from the
control-plane API: registered endpoints are dispatched only by the engine’s
H3/QUIC listener, TCP :443 serves the WebSocket server, no HTTP/2 port is
configured, and no Node or Bun HTTP client speaks HTTP/3. Using the RPC channel
the control-plane API already holds avoids a new listener and any public
exposure.
Two outages motivated that instrumentation, and both are worth recognising
because neither shows up on the transport screen: a cached track whose upstream
had gone away kept accepting subscribes while delivering nothing, and
cross-shard delivery moved zero objects while every log line read ALLOW. In
both cases QUIC connections, RTT, and byte counters can look entirely healthy.
Related
- Telemetry — the gateway’s own metric families, traces, and CDRs
- Telephony Metrics — definitions for the call-level numbers
- Authentication — how a client gets onto the relay in the first place