# Supported providers

> Every speech, language and speech-to-speech vendor a voice agent can use, with the exact provider id for each stage and which ones stream.

Every vendor below is wired into the runtime today. The **id** column is the
literal string you put in `asr_provider`, `llm_provider`, `tts_provider`, or a
`REALTIME` node's `provider` — not a display name.

You do not rebuild anything to change a vendor. You change one config field and
add that vendor's key in the console. Keys are per-tenant, sealed at rest, and
resolved at call setup; they are never written into `agent-config.yml`.

> **NOTE:**
> Several ids accept aliases (`azure` / `azure-speech` / `microsoft` all select
> Azure). The aliases exist so a config written from memory still works. Prefer
> the first id in each row — it is the one the cache keys on.

## Speech-to-text

Set `asr_provider`, and optionally `asr_model` — the runtime forwards it as
the vendor's model hint.

**Streaming** providers hold a socket open and return interim transcripts, so
the agent can barge in mid-sentence. This is what you want on a phone call.

| id | Vendor | Notes |
| --- | --- | --- |
| `deepgram` *(default)* | Deepgram | Interim + final transcripts. Default model `nova-2`. The recommended default for English telephony. |
| `assemblyai` | AssemblyAI | Universal-Streaming v3 — accurate, low latency. |
| `speechmatics` | Speechmatics | Ursa realtime. Strong on enterprise accents and vertical vocabulary. |
| `soniox` | Soniox | Token-level finality and its own endpoint detection. |
| `gladia` | Gladia | Solaria live. Strong multilingual and code-switching. |
| `gradium` | Gradium | Streams native 8 kHz telephony audio — no upsampling on the way in. |
| `sarvam` | Sarvam.ai | Saarika — Indic languages. One key also covers Bulbul TTS. |
| `azure` | Azure Speech | Continuous Recognition. Needs a **region** alongside the key; the region is part of the URL. |
| `smallest` | smallest.ai | Pulse STT, 36 languages. Same bearer token as Lightning TTS. |
| `cartesia` | Cartesia | Ink streaming STT. |
| `elevenlabs` | ElevenLabs | Streaming ASR surface, sharing your ElevenLabs key. |
| `xai` | xAI | Grok STT. Same key as Grok TTS and Grok Voice. |
| `openai-realtime-stt` | OpenAI | Transcription-only use of the Realtime socket (`gpt-4o-transcribe`). |
| `local-whisper` | self-hosted | Whisper on your own `mod_whisper` endpoint. No vendor key. |

**Batch** providers transcribe an utterance after it ends. They cost less and
often read more accurately, but the caller waits for the whole turn before the
agent starts thinking — so use them for offline or non-conversational work,
not for a live call you want to feel fast.

| id | Vendor | Notes |
| --- | --- | --- |
| `groq` | Groq | Whisper large v3 turbo, very fast for a batch pass. |
| `openai` | OpenAI | `gpt-4o-transcribe`. Reuses your OpenAI key. |
| `mistral` | Mistral | Voxtral. Alias `voxtral`. |
| `together` | Together AI | Whisper on Together's endpoint. |
| `fal` | fal.ai | Wizper (Whisper v3 large, serverless). Alias `wizper`. |
| `runpod` | self-hosted | Any OpenAI-compatible transcription endpoint you host. Aliases `selfhost`, `runpod-whisper`. |

## Language models

Set `llm_provider` and `llm_model`. Tool calling runs in this stage — see
[Tool calling](/modalities/voice/runtime/tool-calling).

| id | Vendor | Notes |
| --- | --- | --- |
| `openai` *(default)* | OpenAI | GPT-4o family. |
| `anthropic` | Anthropic | Claude family. Honours in-flight cancellation, so a barge-in stops the running request and you do not pay for a discarded turn. |
| `gemini` | Google | Gemini via the AI Studio API. |
| `groq` | Groq | Llama family at very high tokens/sec. |
| `cerebras` | Cerebras | Wafer-scale Llama inference. |
| `mistral` | Mistral | Mistral family. |
| `together` | Together AI | Open models — Llama, Qwen, DeepSeek. |
| `fireworks` | Fireworks AI | Open-model inference. |
| `openrouter` | OpenRouter | One key that routes to most major vendors. |
| `deepseek` | DeepSeek | Chat and reasoner models. |
| `perplexity` | Perplexity | Sonar online models, web-grounded. |
| `xai` | xAI | Grok. Alias `grok`. |
| `ollama` | self-hosted | A local Ollama endpoint. No API key — just a host:port. |
| `mod_inference` | self-hosted | Our in-process inference gateway. Aliases `clutch`, `local-llm`. |
| `openai-compatible` | anything | Any endpoint speaking the OpenAI chat API. Alias `openai-compat`. |

> **TIP:**
> `llm_fallback_provider` runs a **different vendor** when the primary fails
> after `llm_max_retries`. A vendor incident then degrades one turn instead of
> dropping the call. It is a resilience path, not a load balancer.

## Text-to-speech

Set `tts_provider` and put the vendor's voice id in `tts_voice`.

Streaming providers emit audio as the model's text arrives, which is what keeps
first-audio low. If a streaming session cannot open — a bad key, a missing
voice — the runtime falls back to one-shot synthesis for that turn so the
caller still hears a reply.

| id | Vendor | Streams | Notes |
| --- | --- | --- | --- |
| `elevenlabs` *(default)* | ElevenLabs | yes | `tts_voice` = ElevenLabs `voice_id`. |
| `cartesia` | Cartesia | yes | Sonic over WebSocket. `tts_voice` = voice UUID. Use `cartesia-http` for the one-shot fallback. |
| `rime` | Rime | yes | Mist / Arcana conversational voices. Survives barge-in through the protocol's own clear op. |
| `lmnt` | LMNT | yes | Aurora. Barge-in reopens the session. |
| `smallest` | smallest.ai | yes | Lightning v3.1, multilingual, very low time-to-first-audio. Aliases `smallest-ai`, `lightning`. |
| `deepgram` | Deepgram | yes | Aura voices, e.g. `aura-asteria-en`. |
| `sarvam` | Sarvam.ai | yes | Bulbul — Indic voices. |
| `azure` | Azure Speech | yes | Region required alongside the key. |
| `xai` | xAI | yes | Grok TTS. 24 kHz, downsampled to the call's rate. |
| `openai` | OpenAI | no | One-shot per utterance, reusing your OpenAI key. |
| `local-kokoro` | self-hosted | no | Kokoro through `mod_tts` on loopback. No key. Aliases `kokoro`, `mod_tts`. |
| `runpod` | self-hosted | no | Your own OpenAI-compatible speech endpoint. Aliases `selfhost`, `runpod-kokoro`. |

> **NOTE:**
> **PlayHT** is available for voice **cloning** in Settings → Voices, but is not
> wired as a runtime TTS provider. Clone there, synthesize with one of the
> providers above.

## Speech-to-speech

One duplex model listens and speaks on the same connection — no separate
transcriber, model and synthesizer to chain. Replies start sooner and
interruptions feel instant. These are configured as a `REALTIME` node's
`provider` rather than through the three cascade fields.

| id | Vendor | Notes |
| --- | --- | --- |
| `openai` | OpenAI | Realtime. Server-side VAD, so you add no local turn detector. See [OpenAI Realtime](/modalities/voice/integrations/openai-realtime). |
| `gemini` | Google | Gemini Live. See [Gemini Live](/modalities/voice/integrations/gemini-live). |
| `xai` | xAI | Grok Voice. Aliases `grok`, `grok-voice`. Defaults to `grok-voice-latest` and the `eve` voice. |
| `elevenlabs` | ElevenLabs | Conversational AI. |
| `qwen-omni` | self-hosted | Qwen-Omni over vLLM-Omni, on your own hardware — reachable over TLS WebSocket or a duplex QUIC stream to the sidecar. Aliases `qwen`, `omni`, `vllm-omni`. |

For the shape of a `REALTIME` agent and how to choose duplex over a cascade,
see [BYO speech-to-speech](/modalities/voice/runtime/byo-speech-to-speech).

## Adding your keys

Provider keys live in the console at `agent.telequick.dev` under
**Settings → Providers**. Each vendor takes an API key; Azure additionally
takes a region, and Ollama takes an endpoint instead of a key because there is
nothing to authenticate.

The control plane resolves whichever keys the agent's config actually needs at
call setup. A key you have not set only matters if a stage names its provider.

## Related

- [Bring your own ASR, LLM and TTS](/modalities/voice/runtime/byo-asr-llm-tts) — the cascade in full
- [BYO speech-to-speech](/modalities/voice/runtime/byo-speech-to-speech) — the duplex path
- [Self-hosted inference](/modalities/voice/runtime/self-hosted-inference) — running the models yourself
