Skip to main content
TurnCall is a modular monolith (FastAPI) that orchestrates real-time voice AI agents over phone calls, browser WebRTC, WhatsApp, and SMS/chat. All Pipecat code is isolated in orchestrator/; every other module is framework-agnostic.

System Overview

Inbound Call Flow (Twilio)

Dynamic routing: if the phone number’s routing_target_type is webhook, TurnCall POSTs a call-init request to your server first and applies the returned agent / variables / knowledge context before the pipeline starts. See Pre-Call Init.

Other entry points

POST /v1/calls/outbound creates the Call record and initiates the Twilio call → Twilio hits /webhooks/twilio/voice/outbound → the handler resolves the agent from the Call record by CallSid → same pipeline as inbound.
POST /v1/webrtc/connect with an SDP offer → SmallWebRTCRequestHandler creates the connection and returns the SDP answer → PATCH /v1/webrtc/connect trickles ICE candidates → audio flows peer-to-peer at 16kHz into the same pipeline.
Meta POSTs /webhooks/whatsapp (field calls) → signature validated → Pipecat WhatsAppClient handles the WebRTC SDP exchange → 16kHz SmallWebRTCTransport pipeline.
Inbound text → resolve session (24h TTL) → build LLM history → chat completion → reply. No Pipecat pipeline — it’s a text path through services/.

Real-time Pipeline

Two pipeline modes, selected per agent via pipeline_mode.

Cascade (default, ~800ms processing)

Dashed stages are optional: VoicemailDetector (outbound), KnowledgeRetrieval (auto-mode RAG), video avatar (HeyGen/Tavus, WebRTC + cascade only). Transcript taps sit after STT and after the LLM to record both sides.

Speech-to-Speech (~300ms processing)

A single model handles STT + reasoning + TTS natively over one WebSocket, so the stt/llm/tts config fields are ignored.
Twilio media is 8kHz μ-law on the wire; the serializer converts to/from PCM16. S2S models run at 24kHz, so an internal resampler bridges the rates.

Latency

Two different numbers get called “latency”, and conflating them is the most common way to be surprised by how a agent sounds on a real call. Processing is the sum of each stage’s time-to-first-byte: how long speech-to-text takes to return a transcript, the model to emit its first token, and the voice to produce its first audio chunk. On cascade that is roughly 800 ms; on S2S, where one model does all three, roughly 300 ms. Caller-perceived turnaround is what someone on the phone actually experiences — the gap between their last syllable and your agent’s first. That is processing plus the silence the pipeline must observe before it will believe the caller has stopped talking:
silence_timeout_ms defaults to 800 ms and becomes the VAD stop window. It is the single largest lever you have, and it is a trade: lower it and replies come faster but the agent is likelier to interrupt someone who merely paused; raise it and the agent is more patient but feels sluggish.
Pipecat’s own recommended default for this window is 200 ms, and it will log a warning on every call when yours differs. TurnCall ships 800 ms deliberately — it is more forgiving of hesitant callers on noisy phone lines — but if your callers speak cleanly, lowering silence_timeout_ms to 200–300 ms is the cheapest latency win available to you.
Both figures are engineering estimates for the default provider mix, not a benchmark, and they move with your choice of STT, model and voice. Measure your own: the pipeline emits turn.user_bot_latency_seconds per turn, which is the turnaround number, and per-service ttfb metrics, which sum to the processing number. See Observability.

Call State Machine

Module Responsibilities

Data Model