Provider Matrix
STT Configuration
STT Options
On Deepgram,
extra reaches the full transcription surface: profanity_filter,
diarize, redact, version, utterance_end_ms.
Keyterms
Words a general model mishears — product names, SKUs, surnames — go instt.keyterms as one list. TurnCall maps it to whatever the chosen provider
calls it.
Blanks and duplicates are dropped when the agent is created.
Transcripts are not profanity-filtered by default. Deepgram’s filter
rewrites the words it matches rather than tagging them, so one false positive
silently alters a transcript you store, analyse and receive in
call.ended.
Set "extra": {"profanity_filter": true} if you want it on.LLM Configuration
OpenRouter routes through openrouter.ai with automatic failover — if the primary model rate-limits or errors mid-call, it falls over to the next model in
fallback_models. Voice only (WebRTC / Twilio / WhatsApp voice); not supported on the SMS/Chat text path. The model that answered each turn is recorded on transcript.final events.LLM Options
TTS Configuration
TTS Options
Each provider falls back to its own model and voice. Earlier versions leaked
a Deepgram voice name into the other three, so an agent that chose Cartesia,
OpenAI or ElevenLabs without naming a model and voice was misconfigured.
Pronunciations
Words a voice says wrong — drug names, surnames, SKUs — go intts.pronunciations as one flat word → IPA map. One IPA spelling works
everywhere, because each provider’s markup is generated from it. The outbound
mirror of stt.keyterms.
The value must be real IPA.
{"Metformin": "met-FOR-min"} is a 422 naming
the word: pipecat’s IPA helpers normalize and tokenize but never reject, so a
respelling would be accepted, inserted into the markup, and change nothing you
can hear. /…/ and […] wrappers are stripped and the value stored
canonically, so two spellings of one pronunciation are one config. Blanks are
dropped, and an empty map is the same as not setting one.
ElevenLabs needs a phoneme-capable model. Pipecat reads
<phoneme> tags only
on the models in its own ELEVENLABS_PHONEME_MODELS set, and only with
enable_ssml_parsing — a constructor argument no tts.extra can reach.
TurnCall passes it for exactly the agents that have pronunciations and such a
model, never for everyone: SSML parsing changes how all of that agent’s text
is read. Your model is never downgraded for you.<phoneme> tag reaching a service that is not parsing SSML is
read out loud, measured at 3.5–4.3× the audio. One unusable word is spoken as
written; it never silences the reply.
Whether the built voice can read the markup depends on the model, which an author
changes after the pronunciations are written — so it is a warning, never a
refusal. That, and a “pronunciation” that is the word typed back at us, come
back as warnings on an agent read.
AWS Bedrock
Bedrock is a gateway, not a vendor:provider: "bedrock" with
model: "anthropic.claude-..." names two different companies, and the same
Claude model is reachable through either anthropic or bedrock with
different credentials.
Model ids pass through verbatim in any of the three forms AWS accepts:
Credentials
Bedrock and Nova Sonic share oneaws block on the agent config, so an agent’s
LLM and voice leg always authenticate as the same principal. Credentials are
resolved by the server, in this order:
SSO is a workstation mechanism.
aws sso login is an interactive browser
flow a server cannot perform, and its cached token expires. It works for local
development via aws.profile or the ambient chain; in production use an
instance profile, IRSA, or role_arn.adr/0016-bedrock-and-nova-sonic.md
for why credentials are resolved this way.
API keys
Every provider key resolves in one order, everywhere a project’s work runs — voice calls, chat/SMS/WhatsApp turns, post-call analysis and takeaways, the transfer briefing, knowledge-base embeddings and the eval judge:- The agent’s own key —
llm.api_key(or the agent’sawsblock). - The project’s credential — see Project credentials.
- The platform environment — the variables below.
Platform environment
| TypeSafe (Jev eval judge) |
TYPESAFE_API_KEY |
Ollama requires no API key — just install it locally and run
ollama pull <model>.
custom_openai sends only the agent’s own llm.api_key (falling back to
OPENAI_API_KEY); a project’s OpenAI credential is never sent to an endpoint
the agent names itself.Project credentials
A project can hold its own key per provider account — one Deepgram key serves both transcription and voice, one OpenAI key serves STT, LLM, TTS, S2S and the eval judge. Keys are encrypted at rest underCREDENTIAL_ENCRYPTION_KEY
and never returned by any read.
Deleting a credential is idempotent and returns the project to the platform key.
An
aws project credential is not gated by AWS_AGENT_CREDENTIALS_ENABLED
— that flag governs static keys inside agent config only, because those are
the ones stored unencrypted.
GET /v1/providers lists the valid providers per role and, per account, two
separate signals: platform_key (the environment has one) and
project_credential (this project has one). Together they say whether removing
a credential falls back or ends in failed calls.
When a key is missing or refused
Before a call is built, TurnCall checks every provider the agent uses (STT, LLM and TTS, or the S2S provider, plus the avatar) against all three levels. If one has no key anywhere, the call is refused before it rings through: it endsfailed / pipeline_error with an error.raised event carrying
missing_keys, and POST /v1/webrtc/connect answers
500 pipeline_start_failed with details.missing_keys. AWS is not checked,
since its ambient credential chain is not a key TurnCall can see. Neither are
ollama and custom_openai, which have no account.
A key that exists but is wrong is only found out when the provider refuses it,
partway into the call. That call ends the same way, with provider and
key_rejected: true on the event. See
When a call fails.