silence_timeout_ms (800ms by default), which applies on both pipeline modes.
Configuration
Setpipeline_mode: "s2s" in the agent config. The stt, llm, and tts fields are ignored in S2S mode.
Options
Voices
- OpenAI Realtime — one of a fixed set:
alloy,ash,ballad,coral,echo,sage,shimmer,verse. An unknown voice is rejected on create. - Gemini Live — any Gemini-supported voice (e.g.
Aoede,Charon,Kore,Puck,Zephyr, and more). The set grows per model, so TurnCall doesn’t allowlist it — Gemini validates the voice when the session connects. - Nova Sonic —
matthew(default),tiffany,amy,lupe,carlos. AWS validates on connect.
OpenAI GPT-Live-1
provider: "openai_live" runs OpenAI’s gpt-live-1, a full-duplex model: it
listens and speaks at the same time, and decides for itself when to answer and
when to stop on being talked over. That is the difference from the Realtime API,
which is turn-based.
s2s.model unset and it resolves to gpt-live-1. Set it explicitly and
the value is sent as given.
The agent speaks first: the session opens by asking gpt-live-1 to say the
agent’s first_message, or to greet the caller when there is none. It
delivers the line in its own words — usually verbatim, but not always.
Delegation
gpt-live-1 calls no tools of its own: it listens and speaks, and hands anything that needs a tool or deeper reasoning to a second model — the delegate — while the conversation keeps going, then voices what came back. Without a delegation, anopenai_live agent’s tools are never used.
There are two modes, and an agent with tools gets one without asking:
client— the delegate is your agent’s ownllm: any provider TurnCall supports, on its usual keys. It is the default for anopenai_liveagent with tools (static, MCP, or atool-mode knowledge base) and nodelegationblock. An agent with no tools gets no delegation at all.responses— OpenAI hosts the delegate, the model you name.
query_knowledge for a
tool-mode knowledge base.
In
client mode the delegate’s provider needs a key like any other: one with no
key at any level fails the call before it starts, naming it in missing_keys.
client mode also honours llm.failover: the delegate is an ordinary LLM, so
when it fails in a way it cannot recover from (a rejected key, an unknown
model) the delegation moves to the next backup, recorded as provider.failover
and in failovers on call.ended — as on cascade. Backups need keys too.
llm.failover is refused (422) on every other speech-to-speech agent,
responses mode included.
The agent’s system_prompt reaches both models. TurnCall also appends to
gpt-live-1’s instructions a paragraph telling it to delegate whenever the caller
wants something done, listing each tool’s name and description — so an existing
prompt needs no rewrite.
end_call and do_not_call work as everywhere else, in both modes: the call
ends after gpt-live-1’s goodbye. gpt-live-1 decides by itself whether to
delegate, and it sometimes says goodbye without handing over end_call — the
idle guard then ends the call after two silences.
Refused with a 422: delegation on any other provider (they call tools
themselves), responses without model, and handoff_to_agent among an
openai_live agent’s tools — a handoff would swap the tools but not the live
session’s voice and prompt.
The older s2s.extra.backend_model (with extra.service_tier) keeps working,
read as {"mode": "responses", "model": <it>}.
Differences from openai
pipecat_vad is rejected for this provider: client-side turn detection would
fight a model that is doing its own.
Amazon Nova Sonic
Nova Sonic runs on AWS and shares the agent’saws credential block with the Bedrock LLM provider — same account, same region resolution, no separate key.
Nova Sonic 2 (amazon.nova-2-sonic-v1:0) is the default; the older amazon.nova-sonic-v1:0 still works.
endpointing_sensitivity tunes how quickly Nova Sonic decides the caller has stopped speaking. It lives in extra because it doesn’t fit the server_vad / pipecat_vad split that turn_detection uses.
Sessions roll over about every 6 minutes. Nova Sonic caps session length and TurnCall transparently continues the conversation in a new one — ordinary phone calls hit this routinely. AWS credentials are re-resolved at each rollover, so calls outlasting a temporary credential’s lifetime (assume-role and IRSA are typically ~1 hour) keep working instead of dying mid-sentence.
Gateways and third-party models
Theopenai provider speaks the OpenAI-Realtime WebSocket protocol — and so do OpenAI-Realtime-compatible gateways (Vercel AI Gateway, LiteLLM) and xAI direct. Set s2s.base_url to route the realtime connection through one; provider-prefixed models like xai/grok-voice-think-fast-1.0 or openai/gpt-realtime-2 then flow over the same protocol — no new provider, same pipeline.
Two requirements:
1
Allowlist the gateway URL
base_url is an outbound target, so it’s gated by the same SSRF guard as custom LLM endpoints. Add its wss:// pattern to BYOM_ALLOWED_URL_PATTERNS, e.g. ["wss://ai-gateway.vercel.sh/*"]. A base_url outside the allowlist is rejected before the call starts.2
Use the gateway key
Set
OPENAI_API_KEY to the gateway’s API key — it’s sent as Authorization: Bearer … on the WebSocket.When
base_url is set, the OpenAI voice allowlist is bypassed — the gateway routes to models with their own voice sets (e.g. Grok), which it validates upstream. Pass the target model’s own voice name.Knowledge bases
Two of the three retrieval modes work on a speech-to-speech agent, on every provider:prompt— the document text leads the instructions, ahead ofsystem_prompt. Onopenai_livethe delegate gets it too, unlessdelegation.instructionsreplaces the delegate’s prompt.tool—query_knowledgeis one more tool. Realtime, Gemini Live and Nova Sonic call it themselves; onopenai_livethe delegate does, and gpt-live-1’s delegation paragraph names it. It counts as a tool, so anopenai_liveagent whose only tool it is still gets aclientdelegate.llm.failoveris checked when the agent is saved, before any knowledge base is linked, so such an agent sets"delegation": {"mode": "client"}to be allowed backups.
auto is not supported: it retrieves just before each LLM turn, and a
speech-to-speech model has no such point. The KB is skipped with a warning when
the call is built, and the agent read carries a knowledge_mode_unsupported
entry in warnings naming the mode and provider. Link it in tool mode instead.