# QUIC transport metrics

> How the observability screen derives per-tenant QUIC connection, throughput, RTT, and path-migration numbers from 1-minute ClickHouse rollups.

The observability screen in the TeleQuick portal (`/app/observability`)
is a **layer-4 view of the QUIC transport**, scoped to one tenant. It answers
four questions: how many connections are open, how much data is moving, what
the round-trip latency looks like, and how often connections survived a
network change.

It deliberately shows nothing else. Voice latency, robotics frame rates, and
games ping are per-modality concerns and live on each vertical's own portal
(the agent, robotics, and games subdomains). This page explains what the
numbers on the transport screen mean and how they are computed, because the
tiles and the charts do not all use the same aggregation.

## What the screen covers

All data comes from a single procedure, `observability.quicTransport`, which
takes the org id and a window length in minutes and returns two things:

- `kpi` — the six numbers rendered as tiles across the top.
- `series` — one row per 1-minute bucket, rendered as the four charts.

The procedure is an org-scoped procedure: the tenant id is derived from the
active org and applied as a `tenant_id` filter on every query, so a reader
only ever sees their own transport traffic.

## What a connection (subject) is

The **Connections** number is `countDistinct(subject_id)` within a bucket, not
a count of sessions opened or an event count.

A *subject* is the transport-level identity of one QUIC connection as the
relay records it. Every client that terminates a QUIC connection against a
relay contributes one subject, regardless of which modality it belongs to. In
practice the count mixes:

- voice calls,
- robot links,
- spectator viewers.

Because the count is *distinct subjects per bucket*, a connection that stays
open across ten buckets is counted once in each of those ten buckets — the
series is a concurrency curve, not a cumulative total. A connection that opens
and closes inside the same minute still contributes 1 to that bucket.

The charts and tiles do not break the count down by modality. If you need to
know which vertical the connections belong to, use the per-modality portals
described below.

## How the 1-minute buckets are built

Two ClickHouse rollups back the screen. Both are pre-bucketed at one-minute
granularity and both are filtered to
`bucket > now() - INTERVAL <window> MINUTE`:

| Rollup            | Columns read                                               | Aggregation per bucket                         |
| ----------------- | ---------------------------------------------------------- | ---------------------------------------------- |
| `quic_cwnd_1m_q`  | `subject_id`, `bytes_in`, `bytes_out`, `migrations`        | `countDistinct(subject_id)`, `sum(bytes_in)`, `sum(bytes_out)`, `sum(migrations)` |
| `quic_rtt_1m_q`   | `rtt_p50`, `rtt_p95`, `rtt_p99`                            | `round(avg(...))` of each percentile column    |

The two result sets are then joined in the API by bucket timestamp: the RTT
rows are indexed by bucket and folded into the connection/throughput rows.
Consequences worth knowing:

- The **bucket list comes from `quic_cwnd_1m_q`**. If a bucket exists in the
  RTT rollup but has no connection/byte row, it does not appear on the screen.
- If a bucket has connection rows but **no matching RTT row, the p50/p95/p99
  for that bucket are reported as `0`**. A dip to zero on the latency chart
  means "no RTT sample for that minute", not "instant delivery".
- Buckets are emitted only when there is traffic, so gaps in the charts are
  gaps in data rather than interpolated values.

Because each bucket is exactly 60 seconds, the client converts bytes-per-bucket
into a rate by dividing by 60 (and then by 1024 for the KB/s axis). That
conversion is rounded to a whole number of KB/s, so a very low-rate tenant can
legitimately render as `0 KB/s` on the throughput chart while the bytes series
in the fourth chart is clearly non-zero.

## RTT percentiles and how they are averaged

The percentiles are **not** recomputed from raw samples at query time. They are
already stored per connection-level rollup row as `rtt_p50`, `rtt_p95`, and
`rtt_p99`; the query takes the *arithmetic mean of those percentiles* across all
rows in the bucket and rounds to the nearest millisecond.

That means the charted "p99" is an average of per-connection p99s, not the 99th
percentile of the tenant's whole population of RTT samples. It is the right
shape for spotting a trend or a regression across the fleet, and it is
deliberately insensitive to a single pathological connection: one bad link
drags the averaged p99 up only in proportion to how many connections are in the
bucket.

The subtitle on the chart reads "p50 / p95 / p99 across all connections (ms)"
for this reason — all three lines are fleet-wide roll-ups.

The two RTT tiles are colour-coded so you can triage at a glance. The console
uses these cut-offs for the tile colour only; they are display bands, not a
service commitment:

| Tile     | Green      | Amber          | Red        |
| -------- | ---------- | -------------- | ---------- |
| RTT p50  | < 50 ms    | 50–149 ms      | ≥ 150 ms   |
| RTT p99  | < 200 ms   | 200–499 ms     | ≥ 500 ms   |

## Path migration: what counts as one

QUIC identifies a connection by its connection ID rather than by the
four-tuple, so a client that changes network path keeps the same connection
instead of reconnecting. Each time that happens, the transport records a
migration, and the rollup's `migrations` column carries the count for the
bucket.

The screen reports migrations two ways:

- the **Path migrations** tile sums `migrations` across every bucket in the
  selected window;
- the fourth chart plots per-bucket `migrations` on the left axis, with raw
  `bytes_in` / `bytes_out` on the right axis as a reference trace.

A migration is therefore evidence that the connection **survived** a network
change, not evidence of a failure. The pattern to expect is a spike when a
population of clients moves between networks — for example a handover between
4G and Wi-Fi. A flat zero across a window with mobile clients is the number
worth questioning, because it suggests migration events are not reaching the
rollup.

## Reading the KPI tiles vs the charts

The six tiles do not all use the same aggregation, and the hint text under each
tile tells you which one it is:

| Tile             | Value                                   | Aggregation                       |
| ---------------- | --------------------------------------- | --------------------------------- |
| Connections      | distinct subjects                       | latest 1-minute bucket            |
| Ingress          | `bytes_in_per_sec`                      | window average                    |
| Egress           | `bytes_out_per_sec`                     | window average                    |
| RTT p50          | `rtt_p50_ms`                            | latest 1-minute bucket            |
| RTT p99          | `rtt_p99_ms`                            | latest 1-minute bucket            |
| Path migrations  | `migrations`                            | sum over the window               |

Throughput is averaged on purpose. It is computed as total bytes over the
window divided by `series.length × 60` seconds, so a 30-minute view shows the
steady-state rate rather than whatever the most recent minute happened to do.
Bursty workloads would otherwise make the tile unreadable.

Connections and RTT come from the **latest bucket** because they are
point-in-time facts: you want to know how many connections are up *now* and
what latency they see *now*, not what the average was half an hour ago.

Two practical consequences:

- Changing the window changes the Ingress/Egress and Path-migration tiles, but
  leaves Connections and the RTT tiles alone (the latest bucket is the same
  regardless of how far back you look).
- The Ingress/Egress tiles will not match the right-hand end of the throughput
  chart. The tile is the window average; the chart point is that single
  bucket's rate. If they diverge sharply, the traffic is bursty.

Note also that the `p95` value is present in the series and drawn on the
latency chart, but there is no p95 tile — read it off the chart.

## Time window and refresh

The window selector offers 15 minutes, 30 minutes, 1 hour, 6 hours, and
24 hours. The default is 30 minutes. The procedure accepts any integer number
of minutes from 1 to 1440, so the selector's options are a subset of what the
API allows.

The chosen window is mirrored into the `?w=<minutes>` query parameter via
`history.replaceState`, which means:

- reloading the page keeps your window,
- and you can deep-link a specific window to a colleague.

An out-of-range or unparseable `w` falls back to 30 minutes.

The query refetches every 30 seconds while the screen is open, and the refresh
button in the header forces an immediate refetch. Since the rollups are
1-minute buckets, a manual refresh mid-minute will usually return the same
latest bucket.

Longer windows do not change the bucket size — a 24-hour view is 1-minute
buckets across 24 hours, so expect a dense series.

## Empty and error states

The procedure returns the same zeroed shape in two different situations:

1. **No traffic in the window** — the rollups return no rows for the tenant, so
   `series` is empty and every `kpi` field is `0` (with `window_minutes` echoed
   back).
2. **The ClickHouse query failed** — the handler catches the error and returns
   the same empty payload rather than surfacing an exception.

The screen renders its empty card when there are no series rows and the
connection count is zero. The copy on that card tells the reader that transport
metrics appear once a client (voice gateway, robot bridge, or spectator viewer)
connects to a relay — normally within about a minute, which is the bucket
granularity.

Because the two cases look identical, "no data" on this screen is not proof
that no connections exist. If you expect traffic and see the empty card, treat
it as inconclusive and cross-check against the modality's own portal.

## Where per-modality metrics live instead

This screen stops at the transport layer. Anything that requires knowing what
the bytes *were* belongs to a modality:

- **Voice** — call latency, answer rates, audio quality: the agent portal, and
  [Telephony Metrics](/glossary/metrics).
- **Robotics** — frame rates and control-loop timing: the robotics portal.
- **Games** — ping and room health: the games portal.

Separately from the transport rollups, the observability router exposes
relay data-plane health for **one modality namespace** at a time. Every
modality rides the same MoQ relay, so there is a single implementation
parameterised by namespace, which keeps the consoles from disagreeing about
what a relay stall means. It returns `mod_relay`'s counters as Prometheus
text; the client derives rates by differencing successive samples.

That call is carried over the engine's RPC service (`relay.metrics`), not over
HTTP. `mod_relay` does expose a `/metrics` route, but it is unreachable from the
control-plane API: registered endpoints are dispatched only by the engine's
H3/QUIC listener, TCP `:443` serves the WebSocket server, no HTTP/2 port is
configured, and no Node or Bun HTTP client speaks HTTP/3. Using the RPC channel
the control-plane API already holds avoids a new listener and any public
exposure.

Two outages motivated that instrumentation, and both are worth recognising
because neither shows up on the transport screen: a cached track whose upstream
had gone away kept accepting subscribes while delivering nothing, and
cross-shard delivery moved zero objects while every log line read `ALLOW`. In
both cases QUIC connections, RTT, and byte counters can look entirely healthy.

## Related

- [Telemetry](/platform/telemetry) — the gateway's own metric families, traces, and CDRs
- [Telephony Metrics](/glossary/metrics) — definitions for the call-level numbers
- [Authentication](/concepts/authentication) — how a client gets onto the relay in the first place
