orchestrator/; every other module is framework-agnostic.
System Overview
Inbound Call Flow (Twilio)
Dynamic routing: if the phone number’s
routing_target_type is webhook, TurnCall POSTs a call-init request to your server first and applies the returned agent / variables / knowledge context before the pipeline starts. See Pre-Call Init.Other entry points
Outbound call
Outbound call
POST /v1/calls/outbound creates the Call record and initiates the Twilio call → Twilio hits /webhooks/twilio/voice/outbound → the handler resolves the agent from the Call record by CallSid → same pipeline as inbound.Browser (WebRTC)
Browser (WebRTC)
POST /v1/webrtc/connect with an SDP offer → SmallWebRTCRequestHandler creates the connection and returns the SDP answer → PATCH /v1/webrtc/connect trickles ICE candidates → audio flows peer-to-peer at 16kHz into the same pipeline.WhatsApp voice
WhatsApp voice
Meta POSTs
/webhooks/whatsapp (field calls) → signature validated → Pipecat WhatsAppClient handles the WebRTC SDP exchange → 16kHz SmallWebRTCTransport pipeline.SMS / Chat (text)
SMS / Chat (text)
Inbound text → resolve session (24h TTL) → build LLM history → chat completion → reply. No Pipecat pipeline — it’s a text path through
services/.Real-time Pipeline
Two pipeline modes, selected per agent viapipeline_mode.
Cascade (default, ~800ms processing)
Dashed stages are optional: VoicemailDetector (outbound), KnowledgeRetrieval (auto-mode RAG), video avatar (HeyGen/Tavus, WebRTC + cascade only). Transcript taps sit after STT and after the LLM to record both sides.Speech-to-Speech (~300ms processing)
A single model handles STT + reasoning + TTS natively over one WebSocket, so thestt/llm/tts config fields are ignored.
Twilio media is 8kHz μ-law on the wire; the serializer converts to/from PCM16. S2S models run at 24kHz, so an internal resampler bridges the rates.
Latency
Two different numbers get called “latency”, and conflating them is the most common way to be surprised by how a agent sounds on a real call. Processing is the sum of each stage’s time-to-first-byte: how long speech-to-text takes to return a transcript, the model to emit its first token, and the voice to produce its first audio chunk. On cascade that is roughly 800 ms; on S2S, where one model does all three, roughly 300 ms. Caller-perceived turnaround is what someone on the phone actually experiences — the gap between their last syllable and your agent’s first. That is processing plus the silence the pipeline must observe before it will believe the caller has stopped talking:silence_timeout_ms defaults to 800 ms and becomes the VAD stop window. It
is the single largest lever you have, and it is a trade: lower it and replies
come faster but the agent is likelier to interrupt someone who merely paused;
raise it and the agent is more patient but feels sluggish.
Both figures are engineering estimates for the default provider mix, not a
benchmark, and they move with your choice of STT, model and voice. Measure your
own: the pipeline emits turn.user_bot_latency_seconds per turn, which is the
turnaround number, and per-service ttfb metrics, which sum to the processing
number. See Observability.