agent_config per agent. The left-hand list is every agent_config row that
belongs to the active org, read directly from the database filtered on org_id
and ordered by created_at. The first row is selected automatically when the
screen loads.
Selecting an agent opens the editor. The editor has three view modes and, inside
the default one, four tabs:
A page-level sub-nav switches between Agents (the pool plus this editor) and
Voices (the org’s cloned-voice manager). Voices live here because they are an
agent asset.
The agent pool and the four editor tabs
The config view is split into four panes — Conversation, Models & Voice, Actions, and Advanced. The split is presentation only. All four panes edit the same pending-edit object over the selected row, and one Save writes them together throughadminAgents.updateConfig. Switching tabs never
discards an unsaved change made on another tab.
Per-agent actions in the pool:
Every write goes through these org-scoped, audited BFF procedures rather than a
browser database client, so an agent write is always attributed and always
scoped to the org you are acting in. Creating an agent is gated by the org’s
entitlement limit; the New Agent control is wrapped in the limit gate and will
tell you when the pool is full.
Create an agent: name rules and templates
The agent name is a dispatch identifier, not a display label, so the modal enforces the charset its helper copy promises:- letters, digits and dashes only,
- must start with a letter or a digit,
- at most 64 characters,
- required.
system_prompt and welcome_message only.
It sets nothing else:
New agents are created with coherent cascaded defaults:
pipeline_mode is
cascaded and turn_detection_type is silero.
What each tab covers
Per-agent credentials are columns on the same row —
openai_api_key,
anthropic_api_key, gemini_api_key, elevenlabs_api_key, deepgram_api_key
and ollama_endpoint. They override whatever the org would otherwise use for
that provider. ollama_endpoint must be an http or https URL.
background_audio_url accepts s3, file, https or http.
Cascaded and realtime pipelines
pipeline_mode decides which half of the Models & Voice tab is meaningful.
cascadedruns three separate stages, so it reads the ASR, LLM and TTS provider/model selections and the TTS voice.realtimeruns a single speech-to-speech model, so it reads the realtime model and realtime voice lists instead.
cascaded is the right shape
for any non-realtime model and defaults to silero; realtime sessions use the
provider’s own server_vad. Changing pipeline_mode in the editor
automatically flips turn_detection_type so the pair stays coherent — if you
want a different detector, set it after changing the mode.
Turn detection fields and their accepted ranges
These ranges are enforced in the browser on Save because HTMLmin/max on a
number input is bypassed by typing into it. Leaving a numeric field empty is
valid: the server’s default applies.
Semantic end-of-turn lives in the JSONB
turn_detection_config column, not in
its own columns. The editor exposes three of its params keys:
eou_enabled, eou_threshold and eou_hold_ms.
Preserving turn_detection_config
turn_detection_config is a strict JSONB block with three top-level