Prerequisites
For the browser path:- Docker (runs Postgres, Redis and the API)
- Node 20+ (for the browser client)
- OpenAI API key (or Ollama for a local LLM)
- Deepgram API key (free at console.deepgram.com)
Setup
1
Clone and install
2
Configure environment
.env with your credentials:For the browser path, two keys are all you need:.env
env.example already has a working local default.
Twilio credentials are not required to boot — leave them blank until
you want phone calls.PLATFORM_API_KEY ships with a dev default (dev-platform-key) that gates
project and API-key creation — the examples read it automatically. Set a
strong unique value in production.3
Start the stack
http://localhost:8090. To iterate on server
code with hot reload instead, stop the turncall container (both bind
:8090) and use make run.4
Seed an agent
5
Talk to it
http://localhost:8090), the tc_... key and the agent id from the
previous step, then click Start Call and talk.Audio goes over WebRTC straight to your own machine — speech → Deepgram →
GPT-4o-mini → Deepgram TTS → speech, in roughly 800 ms of processing.Talk to it
That’s the whole loop, with no telephony involved: speech → Deepgram → GPT-4o-mini → Deepgram TTS → speech, in roughly 800 ms of processing. The pause you’ll hear is longer than that: the agent also waits outsilence_timeout_ms (800 ms by default) before it accepts your turn is over.
See Latency.
Things worth trying before you add a phone number:
- Change
system_prompton the agent and reconnect — the personality is one field. - Swap
tts.providertocartesiaorelevenlabsfor a different voice. - Set
pipeline_mode: "s2s"to route through OpenAI Realtime instead, which drops the processing leg to roughly 300 ms. See Speech-to-Speech. - Lower
silence_timeout_msfrom 800 to around 300 if your callers speak cleanly — it comes straight off the pause before every reply. - Point
llm.provideratollamato run the model locally, so nothing but audio transcription leaves your machine.
Answer a real phone call
Everything above needed no Twilio. To have the agent pick up an actual phone, you now also need a Twilio account, a number (~$1/mo) and a tunnel.1
Expose via ngrok
2
Add the Twilio values to .env
.env
3
Bind the number
4
Call your number
The receptionist answers, understands your intent, and routes accordingly.
What You Just Built
Twilio opens a media-stream WebSocket to/ws/media-stream; the pipeline runs until you hang up, then the call is marked completed.
Manual Setup via API
If you prefer to set things up step by step:Create a project
Project and first-API-key creation are gated by the platform credential — sendX-Platform-Key matching your PLATFORM_API_KEY:
Create an API key
Create an agent
Publish the agent
Bind a phone number
Configure Twilio webhooks
Set your Twilio number’s webhook URLs:- Voice URL:
https://xxxx.ngrok.io/webhooks/twilio/voice/inbound(POST) - Status URL:
https://xxxx.ngrok.io/webhooks/twilio/status(POST)
Call your number
That’s it — call the number and talk to your agent.Environment Variables
See Providers for provider-specific configuration and Video Avatar for avatars.