The observability screen in the TeleQuick portal (/app/observability) is a layer-4 view of the QUIC transport, scoped to one tenant. It answers four questions: how many connections are open, how much data is moving, what the round-trip latency looks like, and how often connections survived a network change. It deliberately shows nothing else. Voice latency, robotics frame rates, and games ping are per-modality concerns and live on each vertical’s own portal (the agent, robotics, and games subdomains). This page explains what the numbers on the transport screen mean and how they are computed, because the tiles and the charts do not all use the same aggregation.

What the screen covers

All data comes from a single procedure, observability.quicTransport, which takes the org id and a window length in minutes and returns two things:
  • kpi — the six numbers rendered as tiles across the top.
  • series — one row per 1-minute bucket, rendered as the four charts.
The procedure is an org-scoped procedure: the tenant id is derived from the active org and applied as a tenant_id filter on every query, so a reader only ever sees their own transport traffic.

What a connection (subject) is

The Connections number is countDistinct(subject_id) within a bucket, not a count of sessions opened or an event count. A subject is the transport-level identity of one QUIC connection as the relay records it. Every client that terminates a QUIC connection against a relay contributes one subject, regardless of which modality it belongs to. In practice the count mixes:
  • voice calls,
  • robot links,
  • spectator viewers.
Because the count is distinct subjects per bucket, a connection that stays open across ten buckets is counted once in each of those ten buckets — the series is a concurrency curve, not a cumulative total. A connection that opens and closes inside the same minute still contributes 1 to that bucket. The charts and tiles do not break the count down by modality. If you need to know which vertical the connections belong to, use the per-modality portals described below.

How the 1-minute buckets are built

Two ClickHouse rollups back the screen. Both are pre-bucketed at one-minute granularity and both are filtered to bucket > now() - INTERVAL <window> MINUTE: The two result sets are then joined in the API by bucket timestamp: the RTT rows are indexed by bucket and folded into the connection/throughput rows. Consequences worth knowing:
  • The bucket list comes from quic_cwnd_1m_q. If a bucket exists in the RTT rollup but has no connection/byte row, it does not appear on the screen.
  • If a bucket has connection rows but no matching RTT row, the p50/p95/p99 for that bucket are reported as 0. A dip to zero on the latency chart means “no RTT sample for that minute”, not “instant delivery”.
  • Buckets are emitted only when there is traffic, so gaps in the charts are gaps in data rather than interpolated values.
Because each bucket is exactly 60 seconds, the client converts bytes-per-bucket into a rate by dividing by 60 (and then by 1024 for the KB/s axis). That conversion is rounded to a whole number of KB/s, so a very low-rate tenant can legitimately render as 0 KB/s on the throughput chart while the bytes series in the fourth chart is clearly non-zero.

RTT percentiles and how they are averaged

The percentiles are not recomputed from raw samples at query time. They are already stored per connection-level rollup row as rtt_p50, rtt_p95, and rtt_p99; the query takes the arithmetic mean of those percentiles across all rows in the bucket and rounds to the nearest millisecond. That means the charted “p99” is an average of per-connection p99s, not the 99th percentile of the tenant’s whole population of RTT samples. It is the right shape for spotting a trend or a regression across the fleet, and it is deliberately insensitive to a single pathological connection: one bad link drags the averaged p99 up only in proportion to how many connections are in the bucket. The subtitle on the chart reads “p50 / p95 / p99 across all connections (ms)” for this reason — all three lines are fleet-wide roll-ups. The two RTT tiles are colour-coded so you can triage at a glance. The console uses these cut-offs for the tile colour only; they are display bands, not a service commitment:

Path migration: what counts as one

QUIC identifies a connection by its connection ID rather than by the four-tuple, so a client that changes network path keeps the same connection instead of reconnecting. Each time that happens, the transport records a migration, and the rollup’s migrations column carries the count for the bucket. The screen reports migrations two ways:
  • the Path migrations tile sums migrations across every bucket in the selected window;
  • the fourth chart plots per-bucket migrations on the left axis, with raw bytes_in / bytes_out on the right axis as a reference trace.
A migration is therefore evidence that the connection survived a network change, not evidence of a failure. The pattern to expect is a spike when a population of clients moves between networks — for example a handover between 4G and Wi-Fi. A flat zero across a window with mobile clients is the number worth questioning, because it suggests migration events are not reaching the rollup.

Reading the KPI tiles vs the charts

The six tiles do not all use the same aggregation, and the hint text under each tile tells you which one it is: Throughput is averaged on purpose. It is computed as total bytes over the window divided by series.length × 60 seconds, so a 30-minute view shows the steady-state rate rather than whatever the most recent minute happened to do. Bursty workloads would otherwise make the tile unreadable. Connections and RTT come from the latest bucket because they are point-in-time facts: you want to know how many connections are up now and what latency they see now, not what the average was half an hour ago. Two practical consequences:
  • Changing the window changes the Ingress/Egress and Path-migration tiles, but leaves Connections and the RTT tiles alone (the latest bucket is the same regardless of how far back you look).
  • The Ingress/Egress tiles will not match the right-hand end of the throughput chart. The tile is the window average; the chart point is that single bucket’s rate. If they diverge sharply, the traffic is bursty.
Note also that the p95 value is present in the series and drawn on the latency chart, but there is no p95 tile — read it off the chart.

Time window and refresh

The window selector offers 15 minutes, 30 minutes, 1 hour, 6 hours, and 24 hours. The default is 30 minutes. The procedure accepts any integer number of minutes from 1 to 1440, so the selector’s options are a subset of what the API allows. The chosen window is mirrored into the ?w=<minutes> query parameter via history.replaceState, which means:
  • reloading the page keeps your window,
  • and you can deep-link a specific window to a colleague.
An out-of-range or unparseable w falls back to 30 minutes. The query refetches every 30 seconds while the screen is open, and the refresh button in the header forces an immediate refetch. Since the rollups are 1-minute buckets, a manual refresh mid-minute will usually return the same latest bucket. Longer windows do not change the bucket size — a 24-hour view is 1-minute buckets across 24 hours, so expect a dense series.

Empty and error states

The procedure returns the same zeroed shape in two different situations:
  1. No traffic in the window — the rollups return no rows for the tenant, so series is empty and every kpi field is 0 (with window_minutes echoed back).
  2. The ClickHouse query failed — the handler catches the error and returns the same empty payload rather than surfacing an exception.
The screen renders its empty card when there are no series rows and the connection count is zero. The copy on that card tells the reader that transport metrics appear once a client (voice gateway, robot bridge, or spectator viewer) connects to a relay — normally within about a minute, which is the bucket granularity. Because the two cases look identical, “no data” on this screen is not proof that no connections exist. If you expect traffic and see the empty card, treat it as inconclusive and cross-check against the modality’s own portal.

Where per-modality metrics live instead

This screen stops at the transport layer. Anything that requires knowing what the bytes were belongs to a modality:
  • Voice — call latency, answer rates, audio quality: the agent portal, and Telephony Metrics.
  • Robotics — frame rates and control-loop timing: the robotics portal.
  • Games — ping and room health: the games portal.
Separately from the transport rollups, the observability router exposes relay data-plane health for one modality namespace at a time. Every modality rides the same MoQ relay, so there is a single implementation parameterised by namespace, which keeps the consoles from disagreeing about what a relay stall means. It returns mod_relay’s counters as Prometheus text; the client derives rates by differencing successive samples. That call is carried over the engine’s RPC service (relay.metrics), not over HTTP. mod_relay does expose a /metrics route, but it is unreachable from the control-plane API: registered endpoints are dispatched only by the engine’s H3/QUIC listener, TCP :443 serves the WebSocket server, no HTTP/2 port is configured, and no Node or Bun HTTP client speaks HTTP/3. Using the RPC channel the control-plane API already holds avoids a new listener and any public exposure. Two outages motivated that instrumentation, and both are worth recognising because neither shows up on the transport screen: a cached track whose upstream had gone away kept accepting subscribes while delivering nothing, and cross-shard delivery moved zero objects while every log line read ALLOW. In both cases QUIC connections, RTT, and byte counters can look entirely healthy.
  • Telemetry — the gateway’s own metric families, traces, and CDRs
  • Telephony Metrics — definitions for the call-level numbers
  • Authentication — how a client gets onto the relay in the first place