— while a chart right next to it
shows healthy data, and that is usually correct rather than broken. This
page explains which store backs each panel so you can interpret what you
are looking at.
Every query on the screen is scoped to { orgId, fromMs, toMs } and
re-runs on a 30 s poll. The window is quantized to 30 s buckets, so the
same relative range (“Last 1 hour”) keeps a stable query key between
renders instead of shifting on every repaint.
The three backing stores behind the monitoring screen
A fourth source acts as a fallback only:
telequick.events_raw,
written by the control-plane API’s calls-topic consumer, supplies call
segment counts when tenant-tagged spans are dark.
Alert rules on the Overview tab are not a store at all — they are a
read-only proxy of SigNoz /api/v1/rules. If SigNoz is unreachable or
not configured, the alerts strip renders empty rather than erroring.
CDR-derived tiles and why span timings are not used
The Overview tiles and the Call quality tab both read fromtelequick.cdrs, the rows that the CDR service flushes on
hangup. They do not read span durations.
Span timings on the engine side are not usable as latency numbers:
- Most engine child spans report 0 ms — they are timeline markers, not measured intervals.
- Root spans cover the entire call lifetime, and parent spans are clamped, so a “P95 latency” computed from them is either meaningless or pure noise.
Peak CPS is shown first on purpose. Carrier channel limits trigger on
simultaneous call starts, so peak is the number that matters for trunk
capacity; averaging over a long window flattens bursty traffic into a
figure that looks safe when it is not.
Every CDR-derived tile returns
null — rendered as — — when the window
contains zero calls. It does not render a zero, because “no calls” and “a
measured zero” are different facts.
Tenant tagging: when Spans falls back to Call segments
The fourth Overview tile is a “is telemetry alive?” indicator rather than a performance metric, and it switches identity depending on what is available. The screen first counts spans insignoz_traces.signoz_index_v3 whose
tenant.id attribute matches the tenant. If that count is above zero,
the tile is labelled Spans and shows it.
If it is zero, the tile relabels itself Call segments and shows a
count from telequick.events_raw for the same window. That table
is populated by the control-plane API’s calls-topic consumer, so it stays
live even when the tracing path is not tagging spans with a tenant. The
point is to avoid a false zero while calls are actually flowing.
If calls, spans, and segments are all zero, the tile shows —.
Practical reading: seeing Call segments instead of Spans means
the trace pipeline is not attributing spans to this tenant in this
window. Calls are still being recorded; the Recent traces table below
will be empty, and so will the waterfall drawer.
Recent traces and the waterfall drawer
The Recent traces table resolves traces in two steps. It first collects the distincttrace_ids that have at least one span tagged with the
tenant — checking both the span-level tenant.id attribute and the
resource-level one — then summarizes every span in those traces.
Per trace row:
Because of the span-timing caveat above, treat the Duration column as a
navigation aid for finding the call you want, not as a latency
measurement.
Search is applied server-side across the whole window, not just over the
rows already loaded. The input is debounced and matches either a
case-insensitive substring of the root-span operation name or an exact
trace id. The status chip (
All / OK / Error) is a purely local view
toggle over the returned set — that is why the table footer can read
“n of m shown”.
Opening a row loads the span tree for that trace. This query is
tenant-gated independently: it first verifies that at least one span in
the trace carries your tenant.id, and returns a not-found error
otherwise, so trace ids cannot be enumerated across tenants. Once the
gate passes, all spans in the trace are returned ordered by timestamp,
each with its parent id, service name, kind, start time, duration, status
code, and error flag.
Export JSON serializes the currently loaded page of traces — the rows
the table fetched, after server-side search — into an OTLP-shaped
document for offline analysis. It is not a full export of the window.
Media QoS columns and windows that predate them
The Call quality tab computes per-minute buckets of average MOS, P95 jitter, and P95 packet loss from the media QoS columns on the CDR row (mos, jitter_ms, packet_loss_pct). These carry the RTP leg’s
measurements for each call.
The query only aggregates rows where at least one of those columns is
populated, and each percentile ignores rows where its own column is
absent. Two consequences:
- A window that predates the media QoS columns’ rollout returns zero rows. The panel empty-states instead of inventing values.
- A window that spans the rollout shows QoS only from the point where the columns exist, while the call-count and duration charts beside it cover the whole window. The two panels are not describing different traffic; one simply has fewer populated rows to work with.
direction,
talk-time P50/P95 from duration, setup latency P50/P95/P99 from
answer_time - starting_time, failures per minute against the same
Q.850 rule as the Failure Rate tile, and a window summary that adds
answered count, average duration, and summed cost.
Host metrics are per engine, not per tenant
The System tab readssignoz_metrics.samples_v4, populated by the
engine’s OpenTelemetry collector host-metrics receiver — load average,
CPU, memory, and filesystem.
These are deliberately not tenant-scoped. Every tenant in the
deployment shares the same engine host, so the load average you see is
the same number every other tenant sees. It describes platform health,
not your traffic. Do not correlate a CPU spike here with your own call
volume without checking the CDR-derived panels, which are the only
tenant-attributed view.
Two further caveats follow from the source being a scraped host:
- Host metrics reflect the engine, so they keep moving even when your tenant has placed no calls in the window.
- After an engine restart, samples lag behind the scrape cycle for a short period, so the very latest points on the System tab can be sparse or missing while the CDR and trace panels are already current.
host.name.
Reading an empty panel correctly
Empty is a real answer on this screen, and which store is empty tells you different things.
When a source is unavailable, the screen prefers a null, a soft empty
state, or a relabelled tile over a fabricated zero. If you need to
distinguish “nothing happened” from “we could not ask”, check for the
telemetry banner first, then look at which of the three stores the panel
in question is fed by.
Related
- Telemetry — the streams the gateway emits and the CDR schema it writes
- Telephony Metrics — definitions and healthy ranges for CPS, ASR, MOS, jitter, and packet loss