The Monitoring & observability screen looks like one dashboard, but the numbers on it come from three different backing stores with three different caveats. A tile can read — while a chart right next to it shows healthy data, and that is usually correct rather than broken. This page explains which store backs each panel so you can interpret what you are looking at. Every query on the screen is scoped to { orgId, fromMs, toMs } and re-runs on a 30 s poll. The window is quantized to 30 s buckets, so the same relative range (“Last 1 hour”) keeps a stable query key between renders instead of shifting on every repaint.

The three backing stores behind the monitoring screen

A fourth source acts as a fallback only: telequick.events_raw, written by the control-plane API’s calls-topic consumer, supplies call segment counts when tenant-tagged spans are dark. Alert rules on the Overview tab are not a store at all — they are a read-only proxy of SigNoz /api/v1/rules. If SigNoz is unreachable or not configured, the alerts strip renders empty rather than erroring.

CDR-derived tiles and why span timings are not used

The Overview tiles and the Call quality tab both read from telequick.cdrs, the rows that the CDR service flushes on hangup. They do not read span durations. Span timings on the engine side are not usable as latency numbers:
  • Most engine child spans report 0 ms — they are timeline markers, not measured intervals.
  • Root spans cover the entire call lifetime, and parent spans are clamped, so a “P95 latency” computed from them is either meaningless or pure noise.
So the tiles are computed from CDR columns instead: Peak CPS is shown first on purpose. Carrier channel limits trigger on simultaneous call starts, so peak is the number that matters for trunk capacity; averaging over a long window flattens bursty traffic into a figure that looks safe when it is not. Every CDR-derived tile returns null — rendered as — — when the window contains zero calls. It does not render a zero, because “no calls” and “a measured zero” are different facts.

Tenant tagging: when Spans falls back to Call segments

The fourth Overview tile is a “is telemetry alive?” indicator rather than a performance metric, and it switches identity depending on what is available. The screen first counts spans in signoz_traces.signoz_index_v3 whose tenant.id attribute matches the tenant. If that count is above zero, the tile is labelled Spans and shows it. If it is zero, the tile relabels itself Call segments and shows a count from telequick.events_raw for the same window. That table is populated by the control-plane API’s calls-topic consumer, so it stays live even when the tracing path is not tagging spans with a tenant. The point is to avoid a false zero while calls are actually flowing. If calls, spans, and segments are all zero, the tile shows —. Practical reading: seeing Call segments instead of Spans means the trace pipeline is not attributing spans to this tenant in this window. Calls are still being recorded; the Recent traces table below will be empty, and so will the waterfall drawer.

Recent traces and the waterfall drawer

The Recent traces table resolves traces in two steps. It first collects the distinct trace_ids that have at least one span tagged with the tenant — checking both the span-level tenant.id attribute and the resource-level one — then summarizes every span in those traces. Per trace row: Because of the span-timing caveat above, treat the Duration column as a navigation aid for finding the call you want, not as a latency measurement. Search is applied server-side across the whole window, not just over the rows already loaded. The input is debounced and matches either a case-insensitive substring of the root-span operation name or an exact trace id. The status chip (All / OK / Error) is a purely local view toggle over the returned set — that is why the table footer can read “n of m shown”. Opening a row loads the span tree for that trace. This query is tenant-gated independently: it first verifies that at least one span in the trace carries your tenant.id, and returns a not-found error otherwise, so trace ids cannot be enumerated across tenants. Once the gate passes, all spans in the trace are returned ordered by timestamp, each with its parent id, service name, kind, start time, duration, status code, and error flag. Export JSON serializes the currently loaded page of traces — the rows the table fetched, after server-side search — into an OTLP-shaped document for offline analysis. It is not a full export of the window.

Media QoS columns and windows that predate them

The Call quality tab computes per-minute buckets of average MOS, P95 jitter, and P95 packet loss from the media QoS columns on the CDR row (mos, jitter_ms, packet_loss_pct). These carry the RTP leg’s measurements for each call. The query only aggregates rows where at least one of those columns is populated, and each percentile ignores rows where its own column is absent. Two consequences:
  • A window that predates the media QoS columns’ rollout returns zero rows. The panel empty-states instead of inventing values.
  • A window that spans the rollout shows QoS only from the point where the columns exist, while the call-count and duration charts beside it cover the whole window. The two panels are not describing different traffic; one simply has fewer populated rows to work with.
Other Call quality panels come from plain CDR columns and are available for any window that has calls: calls per minute split by direction, talk-time P50/P95 from duration, setup latency P50/P95/P99 from answer_time - starting_time, failures per minute against the same Q.850 rule as the Failure Rate tile, and a window summary that adds answered count, average duration, and summed cost.

Host metrics are per engine, not per tenant

The System tab reads signoz_metrics.samples_v4, populated by the engine’s OpenTelemetry collector host-metrics receiver — load average, CPU, memory, and filesystem. These are deliberately not tenant-scoped. Every tenant in the deployment shares the same engine host, so the load average you see is the same number every other tenant sees. It describes platform health, not your traffic. Do not correlate a CPU spike here with your own call volume without checking the CDR-derived panels, which are the only tenant-attributed view. Two further caveats follow from the source being a scraped host:
  • Host metrics reflect the engine, so they keep moving even when your tenant has placed no calls in the window.
  • After an engine restart, samples lag behind the scrape cycle for a short period, so the very latest points on the System tab can be sparse or missing while the CDR and trace panels are already current.
If the deployment is later split into per-tenant engines, these become per-host metrics grouped by host.name.

Reading an empty panel correctly

Empty is a real answer on this screen, and which store is empty tells you different things. When a source is unavailable, the screen prefers a null, a soft empty state, or a relabelled tile over a fabricated zero. If you need to distinguish “nothing happened” from “we could not ask”, check for the telemetry banner first, then look at which of the three stores the panel in question is fed by.
  • Telemetry — the streams the gateway emits and the CDR schema it writes
  • Telephony Metrics — definitions and healthy ranges for CPS, ASR, MOS, jitter, and packet loss