# Alerts

> Which voice signals can back an alert rule, how to scope and threshold one, and how a fired alert links back to a call sid.

Alerting turns the operational data that the TeleQuick gateway already
emits into a notification. Nothing new is measured for an alert: every rule
reads one of the existing observability streams — live metrics, Call Detail
Records, or call events. Which stream a rule reads determines how quickly it
can fire, how it can be scoped, and whether the notification can name a
specific call.

This page covers the parts of an alert rule that come from the platform.
The channels and destinations available to your org are listed in the screen
itself.

## Where alert rules live

Alert rules are configured on the **Alerts** screen in the admin section of
the console. The screen body ships from the shared UI package, so the
contact-centre console and the operator console render the same rule form and
the same rule fields.

The screen is part of the control-plane admin surface. Like every other
control-plane call, it authenticates with a long-lived API key scoped to an
org and a set of capabilities, and the control-plane API rejects a key that
lacks the capability scope for the procedure it calls. See
[Authentication](/concepts/authentication).

## Which signals you can alert on

There are three distinct sources, with different shapes.

**Live metrics.** The gateway exposes a fixed, stable set of metric families
on an OpenMetrics endpoint at `http://<gateway-host>:9091/metrics`. These are
aggregates. They carry label values, not call identifiers.

| Metric                                     | Type      | Labels                                   |
| ------------------------------------------ | --------- | ---------------------------------------- |
| `telequick_calls_total`            | Counter   | `tenant`, `trunk`, `direction`, `result` |
| `telequick_calls_active`           | Gauge     | `tenant`, `trunk`                        |
| `telequick_call_duration_seconds`  | Histogram | `tenant`, `trunk`, `result`              |
| `telequick_rtp_packets_total`      | Counter   | `direction`, `codec`                     |
| `telequick_rtp_loss_ratio`         | Gauge     | `tenant`                                 |
| `telequick_jitter_ms`              | Histogram | `tenant`                                 |
| `telequick_estimated_mos`          | Histogram | `tenant`                                 |
| `telequick_codec_engine_lag_us`    | Histogram | `shard`                                  |
| `telequick_quic_handshake_seconds` | Histogram | (none)                                   |
| `telequick_admin_requests_total`   | Counter   | `method`, `result`                       |

**Call Detail Records.** When a call ends, the gateway pushes one CDR row to
ClickHouse. A CDR is per-call and carries `call_sid`, so a rule that reads
CDRs can point at an individual call. The fields that alerting typically
reads:

| Column             | Type                     | Use in a rule                          |
| ------------------ | ------------------------ | -------------------------------------- |
| `status`           | `LowCardinality(String)` | Clear reason / disposition.            |
| `q850_cause`       | `UInt16`                 | ITU-T Q.850 cause code for the clear.  |
| `duration_seconds` | `UInt32`                 | Short-call and stuck-call conditions.  |
| `estimated_mos`    | `Float32`                | Per-call audio quality estimate.       |
| `jitter_ms`        | `Float32`                | Per-call jitter.                       |
| `packets_lost`     | `UInt32`                 | Compare against `packets_sent`.        |
| `codec`            | `LowCardinality(String)` | Isolate a codec-specific regression.   |
| `agent_id`         | `LowCardinality(String)` | Set when the call routed through an AI agent DAG. |
| `trunk_id`         | `LowCardinality(String)` | Attribute a failure to one carrier.    |

**Call events.** The same fields reach your SDK live as a `CallEvent`, with
the full set arriving on `CHANNEL_HANGUP_COMPLETE`. The data is identical to
the CDR; the CDR is the server-side persisted copy.

> **NOTE:**
> Any of the telemetry streams can be disabled in the gateway's deployment
> config. A rule that reads a stream you have opted out of will never fire.

## Scoping a rule with labels

Scope comes from the label or column set of whichever signal the rule reads —
there is no separate scoping dimension.

- `tenant` / `tenant_id` is available on nearly every signal. It is the
  primary scoping key and the one that correlates metrics with CDRs.
- `trunk` / `trunk_id` scopes a rule to a single carrier.
- `direction` is `outbound` or `inbound`.
- `result` / `status` separates answered calls from failures, so a quality
  rule does not average over calls that never carried audio.
- `codec` is present on `telequick_rtp_packets_total` and on CDRs.
- `shard` is present only on `telequick_codec_engine_lag_us`, and
  `method` only on `telequick_admin_requests_total`.

The `telephony_cdrs` table is partitioned by month and indexed on
`(tenant_id, timestamp_ms)`. Rules that read CDRs are cheapest when they pin
`tenant_id` and a bounded time range.

## Choosing a threshold

The platform does not ship default thresholds — the number in a rule is
yours. What the signal's type dictates is the *shape* of the comparison:

| Signal type | Compare on                                                                     |
| ----------- | ------------------------------------------------------------------------------ |
| Counter     | A rate, or a ratio between two label values (for example `result` failures over total). A raw counter value is monotonic and not meaningful as a threshold. |
| Gauge       | The instantaneous value. `telequick_calls_active` and `telequick_rtp_loss_ratio` are gauges. |
| Histogram   | A quantile or a bucket share. `telequick_jitter_ms`, `telequick_estimated_mos`, and `telequick_call_duration_seconds` are histograms, so alert on a quantile rather than a mean. |
| CDR field   | A per-call predicate, or an aggregate over a time range.                        |

`estimated_mos` is an estimate derived from the E-Model R-factor, not a
measured opinion score. For the industry reference ranges that these metrics
are usually judged against — jitter, packet loss, MOS, PDD, ASR, ACD — see
[Telephony Metrics](/glossary/metrics). Those are published healthy ranges,
not platform-enforced limits.

## Evaluation windows

The stream a rule reads bounds how tight its window can be.

- **Metrics** are pulled by a Prometheus-style scrape. Resolution is bounded
  by your scrape interval, so a window shorter than that interval has nothing
  to evaluate.
- **CDRs** are written when the call ends. Anything computed from
  `duration_seconds`, `q850_cause`, `status`, or the per-call quality columns
  is only visible after hangup. A CDR-backed rule cannot catch a problem
  mid-call.
- **Call events** arrive live over the SDK connection, including the state
  transitions that precede `CHANNEL_HANGUP_COMPLETE`. Use these when the
  condition must be detected while the call is still up.
- **Traces** are emitted per RPC over OTLP and cover RPC parse, JWT
  validation, trunk lookup, dialplan node execution, and SIP/RTP setup.

## From a fired alert to the call sid

Whether a notification can name a call depends on the signal behind it.

- **CDR-backed:** the row carries `call_sid` directly, plus `tenant_id`,
  `trunk_id`, `start_timestamp_ms`, `answer_timestamp_ms`, and
  `end_timestamp_ms`.
- **Metric-backed:** metric series carry no call identifier. Use the rule's
  label values and its evaluation window to query `telephony_cdrs` for the
  matching calls — `tenant_id` plus a time range hits the table's index.
- **Traces:** spans carry `call.sid` when applicable, alongside `tenant.id`,
  `method.id`, `transport` (`quic` / `webtransport` / `webrtc` / `sip`), and
  `service.name` = `telequick-gateway`. Correlate with metrics on the
  `tenant` label.
- **Recording:** the CDR's `recording_url` is empty if the call was not
  recorded.
- **Failures with no call record:** if a call fails before any `CallEvent`
  reaches the SDK, there is nothing to link to. Inspect the SIP exchange
  itself via HEPv3 capture to a Homer / heplify-server instance. HEPv3
  doubles the SIP-side packet rate, so enable it only while debugging.

## Related

- [Voice observability overview](/modalities/voice/observability/overview) — what data exists
- [Telemetry](/platform/telemetry) — the four streams, metric names, CDR schema, HEPv3
- [Telephony Metrics](/glossary/metrics) — definitions, formulas, published healthy ranges
- [Authentication](/concepts/authentication) — API keys and capability scopes for control-plane admin calls
