# The Monitoring console screen

> What each tab of the Monitoring & Observability screen measures, which data source backs each tile, and how the time window, filters and refresh cadence behave.

The **Monitoring & Observability** screen in the TeleQuick console shows
trace and call data for the organisation selected in the tenant switcher. The
header shows the active tenant id and the wall-clock time of the last
successful refresh.

The screen has four tabs — Overview, Call Quality, Agent Latency, System — plus
a time-range picker, an **Export JSON** button and a manual **Refresh** button
that are shared across the page.

If the backing queries fail, the page shows a banner reading *"Telemetry
unreachable — metrics are temporarily unavailable"* with the underlying error
message, and the Overview tiles render as `—` instead of throwing.

## What each tab answers

| Tab | Question it answers | Backing procedure |
| --- | ------------------- | ----------------- |
| Overview | Is the platform healthy right now, and what did recent calls look like as traces? | `monitoring.metrics`, `monitoring.traces` |
| Call Quality | How did calls behave per minute — volume, duration, setup latency, failures, media QoS? | `monitoring.callQuality` |
| Agent Latency | How did the voice-AI legs perform? | `monitoring.agentLatency` |
| System | How is the engine host itself doing? | `monitoring.systemMetrics` |

Every tab is driven by the same time range. Switching tabs does not reset the
window.

## Overview tiles and where they are sourced

The four Overview tiles are **CDR-derived, not span-derived**. This matters
when you compare them against a tracing UI: span-derived latency was misleading
here, because engine child spans are zero-duration timeline markers and root
spans cover the whole call lifetime. The tiles therefore read from the
`cdrs` table, and spans are used only as a "is telemetry alive?" signal.

| Tile | What it is | How it is computed |
| ---- | ---------- | ------------------ |
| **P95 Setup** (ms) | Real call setup latency | 95th percentile of `answer_time - starting_time` across CDRs in the window, over rows where `answer_time` is non-null. |
| **Failure Rate** (%) | Share of calls that did not complete cleanly | Calls where `status != 'completed'`, or where the Q.850 cause is neither `0` nor `16` (normal clearing), divided by all calls in the window. |
| **CPS (peak / avg)** | Origination rate | See [Peak versus average CPS](#peak-versus-average-cps). |
| **Spans** / **Call segments** | Telemetry liveness | See [Why the fourth tile changes name](#why-the-fourth-tile-changes-name). |

A tile renders `—` when there were no calls in the window, when the computed
value is not finite, or when the backing query failed.

### Peak versus average CPS

The CPS tile shows two numbers, `peak / avg`.

- **Peak** buckets call start times into **one-second** intervals, counts the
  calls in each bucket, and takes the maximum across the window.
- **Average** is simply total calls divided by the window length in seconds.

Peak is the operationally interesting number for trunk capacity, because a
carrier's channel limit triggers on simultaneous starts. Averaging over a long
window smooths bursty traffic away entirely, so the two numbers can be very far
apart on a campaign workload. Size trunks against the peak.

### Why the fourth tile changes name

The fourth tile is a telemetry-liveness indicator with a fallback:

- If tenant-tagged spans exist in the window, the tile is labelled **Spans**
  and shows the count of spans carrying your `tenant.id` attribute.
- If that count is zero, the tile silently switches to **Call segments** and
  shows the count of call segments recorded for the tenant over the same
  window from the events stream.

The fallback exists so the tile never reports a false zero while calls are
actually flowing. Seeing "Call segments" instead of "Spans" means calls are
being processed but tenant-tagged tracing is dark for that window — worth
investigating on the tracing path, not on the call path. If both are zero and
there are no CDRs, the tile shows `—`.

## Traces list, filters and the JSON export

The Overview tab lists recent traces for the tenant. A trace is included if
**any** span in it carries your `tenant.id` (as a span attribute or a resource
attribute); the whole trace is then summarised.

| Column | Meaning |
| ------ | ------- |
| Operation | The **root span's** name (the span with no parent), so you see e.g. a call-level operation rather than an arbitrary child span. |
| Duration | The longest span duration in the trace, formatted as µs / ms / s. |
| Spans | Number of spans in the trace. |
| Status | Derived from the maximum OTLP status code across the trace's spans. |
| Timestamp | Start of the earliest span in the trace, in your local time. |

Status codes map to the coloured dot as follows:

| Code | Label | Dot |
| ---- | ----- | --- |
| `0` | OK | success |
| `2` | Error | destructive |
| anything else | Unset | warning |

Traces are ordered newest-first.

### The search box filters only the loaded page

The console requests a fixed page of the most recent traces and then filters
**client-side**. The search box matches a case-insensitive substring against the
operation name and the trace id of the rows already on screen; the status
dropdown gates the same rows (`Error` keeps only status `2`; `OK` keeps
everything that is not status `2`, because OTLP `Unset` is the normal default).

The counter next to "Recent Traces" tells you which mode you are in: it reads
`N of M shown` when a filter is narrowing the loaded page, and `M shown`
otherwise.

Consequence: a trace that exists in your time window but fell outside the loaded
page will not appear, no matter what you type. The `monitoring.traces` procedure
itself accepts an optional `search` argument that filters **server-side across
the entire window** — matching a root-span operation-name substring or an exact
trace id — and accepts a `limit`. If you need a window-wide search, call the
procedure directly rather than relying on the box.

### Export JSON

**Export JSON** downloads the traces currently returned by the query (not the
client-filtered subset) as a file named
`traces-<org-id>-<timestamp>.json`. The payload is:

```json
{
  "exported_at": "2026-05-06T10:14:02.000Z",
  "org_id": "org_...",
  "trace_count": 20,
  "traces": [ /* the trace rows as listed above */ ]
}
```

The shape is the OpenTelemetry-style rows the list renders, which
trace-analysis tooling consumes directly. The button is disabled while the page
is loading or when there are no traces.

## Reading the span waterfall

Clicking a trace row opens a drawer titled with the first 12 characters of the
trace id and loads the full span tree via `monitoring.traceSpans`. This query
runs only while the drawer is open.

Each span returns:

| Field | Notes |
| ----- | ----- |
| `spanId` | |
| `parentSpanId` | `null` for the root span — this is what builds the tree. |
| `name` | Operation name. |
| `serviceName` | From the span's `service.name` resource attribute. |
| `kind` | OTLP span kind. |
| `startTimeUnixNano` | Used to position the bar on the timeline. |
| `durationNano` | Used for the bar width. |
| `statusCode` | Same encoding as the list. |
| `hasError` | Boolean error flag carried on the span. |

Spans are returned in ascending start-time order.

Note the interaction with the Overview tiles: engine child spans are emitted as
timeline markers with effectively zero duration, and the root span spans the
whole call. A waterfall is therefore good for *ordering and causality*, and poor
as a source of latency percentiles — use the Call Quality tab's setup-latency
series for that.

**Tenant isolation.** Before returning any spans, the procedure checks that at
least one span in the trace carries the caller's `tenant.id`. If not, it returns
`NOT_FOUND` ("trace not found for tenant"), so trace ids cannot be enumerated
across tenants. The `traceId` input must be a 32-character hex string.

## System tab: shared engine-host metrics

The System tab reports **engine host health, not per-tenant figures**. Its
source is host metrics collected by the engine's OpenTelemetry collector —
load, CPU, memory and filesystem.

Two consequences to internalise before you read the numbers:

- **They are platform-wide.** Every tenant on a deployment shares the same
  engine host, so two tenants looking at this tab see the same values. A spike
  here is not necessarily caused by your traffic. (If engines are later split
  per tenant, these become per-host metrics and the UI groups by `host.name`.)
- **They lag a restart.** Host metrics are sampled and exported on the
  collector's own cadence, so after an engine restart the tab can sit on stale
  or empty series for roughly half a minute before fresh samples land. An empty
  System tab immediately after a deploy is expected; check it again after the
  next refresh tick.

For tenant-scoped resource questions, use the Call Quality and Agent Latency
tabs, which are scoped by `tenant_id` in the CDR queries.

## Refresh cadence and the time window

The time-range picker at the top right drives **every** tab's query — metrics,
traces, call quality, agent latency and system metrics all receive the same
`fromMs` / `toMs`. The default window is the last hour.

Two behaviours are worth knowing:

- **The window is quantised to 30-second buckets.** When the range end is the
  relative token `now`, the resolved timestamps would otherwise change on every
  render and cause a new query each time. Both bounds are floored to 30-second
  boundaries so the query key stays stable. Your window edge may therefore sit
  up to 30 seconds behind true wall clock.
- **Queries poll on a 30-second interval, and revalidation on window focus is
  off.** Tabbing away and back does not trigger a fetch; the page waits for the
  next tick. Use the **Refresh** button to force an immediate refetch of the
  Overview metrics and traces.

The "refreshed HH:MM:SS" stamp in the header updates whenever the metrics or
traces query returns new data, so it reflects the last successful fetch rather
than the last attempt.

The trace-spans query in the drawer does **not** poll; it fetches when you open
a trace.

## Alerting is administered centrally

There is no alert-rule panel on this screen. Alert rules for TeleQuick
deployments are administered centrally by platform administrators in the
observability backend, not per tenant in the console — tenants do not hold
rule-authoring rights, and the read-only panel that once sat at the bottom of
the Overview tab was permanently empty as a result, so it was removed.

The read-only alerts route remains available for operations tooling; it is
simply not wired into this UI. Likewise, the observability backend's own UI is
administrator-only and is not linked from the tenant console — the **Export
JSON** button exists so that tenants can hand trace data to whoever does have
access.

If you need an alert on a metric you can see here, raise it with your
TeleQuick administrator rather than looking for a button on this page.

## Related

- [Telemetry](/platform/telemetry) — the streams these views are built on, and the CDR schema
- [Telephony Metrics](/glossary/metrics) — definitions and healthy ranges for CPS, setup latency, MOS, jitter and loss
