— instead of throwing.
What each tab answers
Every tab is driven by the same time range. Switching tabs does not reset the
window.
Overview tiles and where they are sourced
The four Overview tiles are CDR-derived, not span-derived. This matters when you compare them against a tracing UI: span-derived latency was misleading here, because engine child spans are zero-duration timeline markers and root spans cover the whole call lifetime. The tiles therefore read from thecdrs table, and spans are used only as a “is telemetry alive?” signal.
A tile renders
— when there were no calls in the window, when the computed
value is not finite, or when the backing query failed.
Peak versus average CPS
The CPS tile shows two numbers,peak / avg.
- Peak buckets call start times into one-second intervals, counts the calls in each bucket, and takes the maximum across the window.
- Average is simply total calls divided by the window length in seconds.
Why the fourth tile changes name
The fourth tile is a telemetry-liveness indicator with a fallback:- If tenant-tagged spans exist in the window, the tile is labelled Spans
and shows the count of spans carrying your
tenant.idattribute. - If that count is zero, the tile silently switches to Call segments and shows the count of call segments recorded for the tenant over the same window from the events stream.
—.
Traces list, filters and the JSON export
The Overview tab lists recent traces for the tenant. A trace is included if any span in it carries yourtenant.id (as a span attribute or a resource
attribute); the whole trace is then summarised.
Status codes map to the coloured dot as follows:
Traces are ordered newest-first.
The search box filters only the loaded page
The console requests a fixed page of the most recent traces and then filters client-side. The search box matches a case-insensitive substring against the operation name and the trace id of the rows already on screen; the status dropdown gates the same rows (Error keeps only status 2; OK keeps
everything that is not status 2, because OTLP Unset is the normal default).
The counter next to “Recent Traces” tells you which mode you are in: it reads
N of M shown when a filter is narrowing the loaded page, and M shown
otherwise.
Consequence: a trace that exists in your time window but fell outside the loaded
page will not appear, no matter what you type. The monitoring.traces procedure
itself accepts an optional search argument that filters server-side across
the entire window — matching a root-span operation-name substring or an exact
trace id — and accepts a limit. If you need a window-wide search, call the
procedure directly rather than relying on the box.
Export JSON
Export JSON downloads the traces currently returned by the query (not the client-filtered subset) as a file namedtraces-<org-id>-<timestamp>.json. The payload is:
Reading the span waterfall
Clicking a trace row opens a drawer titled with the first 12 characters of the trace id and loads the full span tree viamonitoring.traceSpans. This query
runs only while the drawer is open.
Each span returns:
Spans are returned in ascending start-time order.
Note the interaction with the Overview tiles: engine child spans are emitted as
timeline markers with effectively zero duration, and the root span spans the
whole call. A waterfall is therefore good for ordering and causality, and poor
as a source of latency percentiles — use the Call Quality tab’s setup-latency
series for that.
Tenant isolation. Before returning any spans, the procedure checks that at
least one span in the trace carries the caller’s
tenant.id. If not, it returns
NOT_FOUND (“trace not found for tenant”), so trace ids cannot be enumerated
across tenants. The traceId input must be a 32-character hex string.
System tab: shared engine-host metrics
The System tab reports engine host health, not per-tenant figures. Its source is host metrics collected by the engine’s OpenTelemetry collector — load, CPU, memory and filesystem. Two consequences to internalise before you read the numbers:- They are platform-wide. Every tenant on a deployment shares the same
engine host, so two tenants looking at this tab see the same values. A spike
here is not necessarily caused by your traffic. (If engines are later split
per tenant, these become per-host metrics and the UI groups by
host.name.) - They lag a restart. Host metrics are sampled and exported on the collector’s own cadence, so after an engine restart the tab can sit on stale or empty series for roughly half a minute before fresh samples land. An empty System tab immediately after a deploy is expected; check it again after the next refresh tick.
tenant_id in the CDR queries.
Refresh cadence and the time window
The time-range picker at the top right drives every tab’s query — metrics, traces, call quality, agent latency and system metrics all receive the samefromMs / toMs. The default window is the last hour.
Two behaviours are worth knowing:
- The window is quantised to 30-second buckets. When the range end is the
relative token
now, the resolved timestamps would otherwise change on every render and cause a new query each time. Both bounds are floored to 30-second boundaries so the query key stays stable. Your window edge may therefore sit up to 30 seconds behind true wall clock. - Queries poll on a 30-second interval, and revalidation on window focus is off. Tabbing away and back does not trigger a fetch; the page waits for the next tick. Use the Refresh button to force an immediate refetch of the Overview metrics and traces.
Alerting is administered centrally
There is no alert-rule panel on this screen. Alert rules for TeleQuick deployments are administered centrally by platform administrators in the observability backend, not per tenant in the console — tenants do not hold rule-authoring rights, and the read-only panel that once sat at the bottom of the Overview tab was permanently empty as a result, so it was removed. The read-only alerts route remains available for operations tooling; it is simply not wired into this UI. Likewise, the observability backend’s own UI is administrator-only and is not linked from the tenant console — the Export JSON button exists so that tenants can hand trace data to whoever does have access. If you need an alert on a metric you can see here, raise it with your TeleQuick administrator rather than looking for a button on this page.Related
- Telemetry — the streams these views are built on, and the CDR schema
- Telephony Metrics — definitions and healthy ranges for CPS, setup latency, MOS, jitter and loss