Every vendor below is wired into the runtime today. The id column is the literal string you put in asr_provider, llm_provider, tts_provider, or a REALTIME node’s provider — not a display name. You do not rebuild anything to change a vendor. You change one config field and add that vendor’s key in the console. Keys are per-tenant, sealed at rest, and resolved at call setup; they are never written into agent-config.yml.
Several ids accept aliases (azure / azure-speech / microsoft all select Azure). The aliases exist so a config written from memory still works. Prefer the first id in each row — it is the one the cache keys on.

Speech-to-text

Set asr_provider, and optionally asr_model — the runtime forwards it as the vendor’s model hint. Streaming providers hold a socket open and return interim transcripts, so the agent can barge in mid-sentence. This is what you want on a phone call. Batch providers transcribe an utterance after it ends. They cost less and often read more accurately, but the caller waits for the whole turn before the agent starts thinking — so use them for offline or non-conversational work, not for a live call you want to feel fast.

Language models

Set llm_provider and llm_model. Tool calling runs in this stage — see Tool calling.
llm_fallback_provider runs a different vendor when the primary fails after llm_max_retries. A vendor incident then degrades one turn instead of dropping the call. It is a resilience path, not a load balancer.

Text-to-speech

Set tts_provider and put the vendor’s voice id in tts_voice. Streaming providers emit audio as the model’s text arrives, which is what keeps first-audio low. If a streaming session cannot open — a bad key, a missing voice — the runtime falls back to one-shot synthesis for that turn so the caller still hears a reply.
PlayHT is available for voice cloning in Settings → Voices, but is not wired as a runtime TTS provider. Clone there, synthesize with one of the providers above.

Speech-to-speech

One duplex model listens and speaks on the same connection — no separate transcriber, model and synthesizer to chain. Replies start sooner and interruptions feel instant. These are configured as a REALTIME node’s provider rather than through the three cascade fields. For the shape of a REALTIME agent and how to choose duplex over a cascade, see BYO speech-to-speech.

Adding your keys

Provider keys live in the console at agent.telequick.dev under Settings → Providers. Each vendor takes an API key; Azure additionally takes a region, and Ollama takes an endpoint instead of a key because there is nothing to authenticate. The control plane resolves whichever keys the agent’s config actually needs at call setup. A key you have not set only matters if a stage names its provider.