Skip to main content
Speech-to-Speech (S2S) mode skips separate STT and TTS stages — the model handles audio natively, cutting processing to roughly 300ms. The caller-perceived pause also includes silence_timeout_ms (800ms by default), which applies on both pipeline modes.

Configuration

Set pipeline_mode: "s2s" in the agent config. The stt, llm, and tts fields are ignored in S2S mode.

Options

Voices

  • OpenAI Realtime — one of a fixed set: alloy, ash, ballad, coral, echo, sage, shimmer, verse. An unknown voice is rejected on create.
  • Gemini Live — any Gemini-supported voice (e.g. Aoede, Charon, Kore, Puck, Zephyr, and more). The set grows per model, so TurnCall doesn’t allowlist it — Gemini validates the voice when the session connects.
  • Nova Sonic — matthew (default), tiffany, amy, lupe, carlos. AWS validates on connect.

OpenAI GPT-Live-1

provider: "openai_live" runs OpenAI’s gpt-live-1, a full-duplex model: it listens and speaks at the same time, and decides for itself when to answer and when to stop on being talked over. That is the difference from the Realtime API, which is turn-based.
Leave s2s.model unset and it resolves to gpt-live-1. Set it explicitly and the value is sent as given. The agent speaks first: the session opens by asking gpt-live-1 to say the agent’s first_message, or to greet the caller when there is none. It delivers the line in its own words — usually verbatim, but not always.

Delegation

gpt-live-1 calls no tools of its own: it listens and speaks, and hands anything that needs a tool or deeper reasoning to a second model — the delegate — while the conversation keeps going, then voices what came back. Without a delegation, an openai_live agent’s tools are never used. There are two modes, and an agent with tools gets one without asking:
  • client — the delegate is your agent’s own llm: any provider TurnCall supports, on its usual keys. It is the default for an openai_live agent with tools (static, MCP, or a tool-mode knowledge base) and no delegation block. An agent with no tools gets no delegation at all.
  • responses — OpenAI hosts the delegate, the model you name.
Either way the delegate runs your agent’s own tools — webhook and MCP alike, mocked in evals and recorded like on any provider — and query_knowledge for a tool-mode knowledge base.
In client mode the delegate’s provider needs a key like any other: one with no key at any level fails the call before it starts, naming it in missing_keys. client mode also honours llm.failover: the delegate is an ordinary LLM, so when it fails in a way it cannot recover from (a rejected key, an unknown model) the delegation moves to the next backup, recorded as provider.failover and in failovers on call.ended — as on cascade. Backups need keys too. llm.failover is refused (422) on every other speech-to-speech agent, responses mode included. The agent’s system_prompt reaches both models. TurnCall also appends to gpt-live-1’s instructions a paragraph telling it to delegate whenever the caller wants something done, listing each tool’s name and description — so an existing prompt needs no rewrite. end_call and do_not_call work as everywhere else, in both modes: the call ends after gpt-live-1’s goodbye. gpt-live-1 decides by itself whether to delegate, and it sometimes says goodbye without handing over end_call — the idle guard then ends the call after two silences. Refused with a 422: delegation on any other provider (they call tools themselves), responses without model, and handoff_to_agent among an openai_live agent’s tools — a handoff would swap the tools but not the live session’s voice and prompt. The older s2s.extra.backend_model (with extra.service_tier) keeps working, read as {"mode": "responses", "model": <it>}.

Differences from openai

pipecat_vad is rejected for this provider: client-side turn detection would fight a model that is doing its own.

Amazon Nova Sonic

Nova Sonic runs on AWS and shares the agent’s aws credential block with the Bedrock LLM provider — same account, same region resolution, no separate key. Nova Sonic 2 (amazon.nova-2-sonic-v1:0) is the default; the older amazon.nova-sonic-v1:0 still works.
endpointing_sensitivity tunes how quickly Nova Sonic decides the caller has stopped speaking. It lives in extra because it doesn’t fit the server_vad / pipecat_vad split that turn_detection uses.
endpointing_sensitivity is Nova Sonic 2 only — it is silently ignored on amazon.nova-sonic-v1:0.
Sessions roll over about every 6 minutes. Nova Sonic caps session length and TurnCall transparently continues the conversation in a new one — ordinary phone calls hit this routinely. AWS credentials are re-resolved at each rollover, so calls outlasting a temporary credential’s lifetime (assume-role and IRSA are typically ~1 hour) keep working instead of dying mid-sentence.

Gateways and third-party models

The openai provider speaks the OpenAI-Realtime WebSocket protocol — and so do OpenAI-Realtime-compatible gateways (Vercel AI Gateway, LiteLLM) and xAI direct. Set s2s.base_url to route the realtime connection through one; provider-prefixed models like xai/grok-voice-think-fast-1.0 or openai/gpt-realtime-2 then flow over the same protocol — no new provider, same pipeline. Two requirements:
1

Allowlist the gateway URL

base_url is an outbound target, so it’s gated by the same SSRF guard as custom LLM endpoints. Add its wss:// pattern to BYOM_ALLOWED_URL_PATTERNS, e.g. ["wss://ai-gateway.vercel.sh/*"]. A base_url outside the allowlist is rejected before the call starts.
2

Use the gateway key

Set OPENAI_API_KEY to the gateway’s API key — it’s sent as Authorization: Bearer … on the WebSocket.
When base_url is set, the OpenAI voice allowlist is bypassed — the gateway routes to models with their own voice sets (e.g. Grok), which it validates upstream. Pass the target model’s own voice name.

Knowledge bases

Two of the three retrieval modes work on a speech-to-speech agent, on every provider:
  • prompt — the document text leads the instructions, ahead of system_prompt. On openai_live the delegate gets it too, unless delegation.instructions replaces the delegate’s prompt.
  • tool — query_knowledge is one more tool. Realtime, Gemini Live and Nova Sonic call it themselves; on openai_live the delegate does, and gpt-live-1’s delegation paragraph names it. It counts as a tool, so an openai_live agent whose only tool it is still gets a client delegate. llm.failover is checked when the agent is saved, before any knowledge base is linked, so such an agent sets "delegation": {"mode": "client"} to be allowed backups.
auto is not supported: it retrieves just before each LLM turn, and a speech-to-speech model has no such point. The KB is skipped with a warning when the call is built, and the agent read carries a knowledge_mode_unsupported entry in warnings naming the mode and provider. Link it in tool mode instead.

Pipeline

Required API Keys

S2S mode cannot be combined with voicemail_detection.enabled: true.