# Infrastructure probes and managed storage

> What the Infrastructure tab probes, how to read each probe status, and how the managed per-tenant recording bucket is provisioned and overridden.

The **Infrastructure** tab in Settings answers one question: can this
deployment reach the services it depends on right now? It does that by
firing a live probe at each service every time you open the tab, plus on
demand via **Re-probe**.

Storage is the other half of the same picture. The Settings **Storage**
panel reports the managed, per-tenant recording bucket that TeleQuick
provisions for your org, and tells you whether recordings will land there
or in your own S3.

## What the Infrastructure tab probes

Six services, split by who can reach them.

| Service        | Probed from | Configured by        |
| -------------- | ----------- | -------------------- |
| Supabase       | Browser     | `SUPABASE_URL`       |
| The platform API | Browser   | `TRPC_URL`           |
| Admin QUIC     | Not probed  | `QUIC_URL`           |
| ClickHouse     | API server  | `CLICKHOUSE_URL`     |
| CGRates        | API server  | `CGRATES_URL`        |
| SigNoz         | API server  | `SIGNOZ_URL`         |
| Homer          | API server  | `HOMER_URL`          |

Each row in the panel also shows the service's *purpose* — portal
database and auth, API gateway to telemetry services, call analytics
database, per-second billing engine, APM metrics and traces, SIP packet
capture — so the panel doubles as a map of what the deployment is made
of.

If a browser-probed service has no URL configured, the panel renders
`(not configured)` in place of the URL.

## What each probe actually calls

The probes are deliberately cheap. None of them mutate anything.

| Service      | Request                                                                   |
| ------------ | ------------------------------------------------------------------------- |
| Supabase     | `HEAD {SUPABASE_URL}/rest/v1/`                                            |
| Platform API | `GET {TRPC_URL}/health` (falls back to `http://localhost:3001` if unset)   |
| Admin QUIC   | none — QUIC is not reachable with `fetch()` from a browser                 |
| ClickHouse   | `GET {CLICKHOUSE_URL}/ping` with HTTP Basic auth built from `CLICKHOUSE_USER` / `CLICKHOUSE_PASSWORD` |
| CGRates      | `POST {CGRATES_URL}` with the JSON-RPC body `{"method":"APIerSv1.Ping"}`   |
| SigNoz       | `GET {SIGNOZ_URL}/api/v1/version`                                         |
| Homer        | `POST {HOMER_URL}/api/v3/user/auth` with `HOMER_ADMIN_USER` / `HOMER_ADMIN_PASSWORD` |

Two of these are worth explaining:

- **CGRates** has no GET health route, so the probe posts the cheapest
  RPC in the API — `APIerSv1.Ping` — which always returns something if
  the engine is up.
- **Homer** exposes its login endpoint rather than a health endpoint. The
  probe posts the configured admin credentials, but a rejection is still
  a useful signal: a `401` or `422` response proves Homer is listening
  and speaking HTTP, so the probe counts it as reachable rather than
  down.

Server-side probes also report `latency_ms` and, when there is something
worth saying, a `detail` string alongside the status.

## Reading ok, unreachable, auth needed, and manual

Every probe resolves to one of four statuses.

| Status         | Rendered as   | Meaning                                                                 |
| -------------- | ------------- | ----------------------------------------------------------------------- |
| `ok`           | `ok`          | The service answered. It is up and routable from the probing side.      |
| `unreachable`  | `unreachable` | The request failed, or returned a status the probe does not treat as reachable. DNS failure, connection refused, TLS failure, or a CORS block on a browser probe all land here. |
| `unauthorized` | `auth needed` | The service answered, but rejected the credentials the probe presented. The network path is fine; the configured username, password, or token is not. |
| `unknown`      | `manual`      | No probe was attempted. Verify this one yourself.                       |

There is also a transient `pending` state while a browser probe is in
flight; the panel shows it as a pulsing amber dot.

Two subtleties:

- **`ok` is a reachability claim, not an authorization claim.** For
  browser probes, a `401` response is deliberately scored as `ok` — the
  browser has no credentials to present, and proving the service is
  listening is the whole point. So a green Supabase row does not mean
  the portal's key is valid.
- **`manual` is the expected steady state for Admin QUIC.** QUIC is a
  binary transport; a browser `fetch()` cannot speak it. The row exists
  to show you the configured `QUIC_URL`, not to health-check it. Confirm
  the admin RPC path by making an actual admin call.

## Browser-side vs server-side probes

Supabase and the platform API are probed **directly from the browser**.
Both are already browser-reachable with CORS in a working deployment, so
probing them client-side tests the exact path the console itself uses. If
the browser probe fails but the service is otherwise healthy, you have
learned something real: the console cannot reach it either.

ClickHouse, CGRates, SigNoz, and Homer are probed **on the API server**,
through `diagnostics.probeAll`. That is not a network-topology decision —
it is a secrets decision. Those probes need the ClickHouse basic-auth
credentials and the Homer admin username and password to produce a
meaningful result. Shipping those to the SPA so it could probe them
itself would put infrastructure credentials in browser memory and in the
network tab of anyone with the console open.

So the rule the panel follows is: **if a useful probe needs a secret, the
probe runs server-side and only the verdict crosses the wire.** The
browser receives a status, a latency, and an optional detail string. It
never receives the credentials, and it never receives the internal URL
those credentials were used against.

`diagnostics.probeAll` takes no input and is not org-scoped — it reports
the health of the deployment, not of your tenant. It does require an
authenticated session.

## Public URLs vs internal URLs

Each back-end service has two addresses in the deployment config, and
they do different jobs.

- `CLICKHOUSE_URL`, `CGRATES_URL`, `SIGNOZ_URL`, `HOMER_URL` — the
  **internal** addresses. The API server dials these to run the probe.
  They frequently resolve only inside the deployment's network (container
  names, private DNS, `localhost`). These are **never returned to the
  browser**.
- `CLICKHOUSE_PUBLIC_URL`, `CGRATES_PUBLIC_URL`, `SIGNOZ_PUBLIC_URL`,
  `HOMER_PUBLIC_URL` — the **public** addresses. These are what the panel
  shows you and what you click through to. They are the only URLs in the
  probe response.

If a service has no `*_PUBLIC_URL` configured, the probe returns an empty
string for `url` and the console renders "no public URL" rather than
offering you a `localhost` link that would be dead in your browser. The
probe still runs and still reports a status — a service can be perfectly
healthy and simply have no externally routable address, which is normal
for ClickHouse and CGRates in a locked-down deployment.

The practical consequence: a row showing `ok` with no public URL is a
correctly configured service you cannot open a UI for. A row showing
`unreachable` with a public URL set is a service the API server cannot
talk to, regardless of whether your browser could.

## Re-probe rate limit

`diagnostics.probeAll` is rate limited to **10 calls per 60 seconds** per
caller, under the `diagnostics.probeAll` bucket. Opening the tab consumes
one call, and each **Re-probe** click consumes one more. Client-side
probes (Supabase, the platform API) are not rate limited, because they
never touch the API server — but the **Re-probe** button fires both
halves together, so hitting it repeatedly will burn the server-side
budget.

The tab is configured to refetch on mount with no cache, so navigating
away and back always produces a fresh probe rather than a stale verdict.

## Managed recording storage

TeleQuick can provision an S3-compatible bucket scoped to your tenant
so call recordings have somewhere to land without you configuring object
storage at all. The Settings **Storage** panel reads this from
`storage.info`, which is scoped to the active org.

When the deployment has managed storage configured, the panel shows:

| Field          | Meaning                                                                 |
| -------------- | ----------------------------------------------------------------------- |
| Bucket         | The bucket recordings are written to.                                   |
| Prefix         | The key prefix reserved for your tenant — your org id followed by `/`.  |
| Access Key     | The access key minted for your tenant.                                  |
| Retention      | How long objects are kept. Managed storage reports **30 days**.         |
| Usage          | Current consumption, as bytes and as an object count.                   |
| Endpoint       | The S3 API endpoint to point a client at.                               |

The endpoint the panel shows is the deployment's public storage URL when
one is configured, and the internal API URL otherwise — so in a
deployment without a public storage address, that value may not resolve
from your machine.

**Provisioning is lazy and idempotent.** If your org was created before
the storage onboarding hook existed, or onboarding ran while the storage
backend was briefly unavailable, credentials are minted the first time
anyone opens the Storage panel. The result is cached, so opening the
panel repeatedly does not churn credentials.

If the deployment has no storage backend configured at all, the panel
reports managed storage as unavailable and tells you plainly: recordings
will only be stored if your trunk carries a custom `recording_url`. In
that state the panel still shows the bucket name and the `{orgId}/`
prefix the deployment *would* use, with no access key and no endpoint.

## Overriding managed storage with a trunk recording_url

Managed storage is the fallback, not the mandate. The precedence is:

1. **Trunk has a `recording_url` set** — recordings go there. That is
   your bucket, your credentials, your lifecycle policy. The managed
   bucket's 30-day retention does not apply to objects it never
   receives, and the Usage figure in the Storage panel will not grow.
2. **Trunk has no `recording_url`** — recordings go to the managed
   bucket under your tenant prefix, and the managed retention applies.

This is the lever to pull when 30 days is not the retention you need.
Configure `recording_url` on the trunk to bring your own S3, and
retention becomes entirely a property of your bucket's lifecycle rules.

Note that the override is **per trunk**, not per org. A deployment with a
mix of trunks can have some recordings in your bucket and some in the
managed one; the Storage panel's usage numbers only ever describe the
managed bucket.

## Recording playback URLs

The console does not embed durable links to recording objects. When you
click play on a row in Recordings, it requests a **presigned download URL
valid for one hour**, generated on demand. URLs are minted lazily — only
for rows you actually listen to — rather than for every row in a listing.

The lookup is tenant-scoped at the query level: the recording's
`recording_url` is read from the CDR row for that call id *and* filtered
to your org's `tenant_id`. Guessing another tenant's `call_sid` does not
produce a playable URL.

## Branding is deployment-wide

The **Branding** tab sits beside Infrastructure in Settings, but it is
not an org profile field, and that distinction matters when you change
it. Branding is a **deployment-wide white-label setting**: it is the one
setting that changes what *every* console in the deployment renders — the
portal, the voice console, and each modality SPA.

Changing the organization name under **Account Information** renames one
org. Changing branding restyles the whole installation for everyone on
it. That is why it has its own tab rather than living next to the org
name field.

## Related

- [Telemetry](/platform/telemetry) — what ClickHouse, SigNoz, and Homer actually store
- [Authentication](/concepts/authentication) — API keys and the org admin token
- [Telephony Metrics](/glossary/metrics) — reading the numbers these services collect
