Skip to main content
Two paths. Start with the browser one — no Twilio account, no phone number, no tunnel, so you find out whether you like this before spending money.

Prerequisites

For the browser path:
  • Docker (runs Postgres, Redis and the API)
  • Node 20+ (for the browser client)
  • OpenAI API key (or Ollama for a local LLM)
  • Deepgram API key (free at console.deepgram.com)
Additionally for phone calls: a Twilio account with a number, and ngrok. Python 3.12+ is only needed for host development.

Setup

1

Clone and install

2

Configure environment

Edit .env with your credentials:For the browser path, two keys are all you need:
.env
Everything else in env.example already has a working local default. Twilio credentials are not required to boot — leave them blank until you want phone calls.PLATFORM_API_KEY ships with a dev default (dev-platform-key) that gates project and API-key creation — the examples read it automatically. Set a strong unique value in production.
3

Start the stack

The API is now serving on http://localhost:8090. To iterate on server code with hot reload instead, stop the turncall container (both bind :8090) and use make run.
4

Seed an agent

Creates a project, an API key and a published receptionist agent, and prints both — the key is shown once. No phone number involved.
5

Talk to it

Open http://localhost:5174, paste in the server URL (http://localhost:8090), the tc_... key and the agent id from the previous step, then click Start Call and talk.Audio goes over WebRTC straight to your own machine — speech → Deepgram → GPT-4o-mini → Deepgram TTS → speech, in roughly 800 ms of processing.

Talk to it

That’s the whole loop, with no telephony involved: speech → Deepgram → GPT-4o-mini → Deepgram TTS → speech, in roughly 800 ms of processing. The pause you’ll hear is longer than that: the agent also waits out silence_timeout_ms (800 ms by default) before it accepts your turn is over. See Latency. Things worth trying before you add a phone number:
  • Change system_prompt on the agent and reconnect — the personality is one field.
  • Swap tts.provider to cartesia or elevenlabs for a different voice.
  • Set pipeline_mode: "s2s" to route through OpenAI Realtime instead, which drops the processing leg to roughly 300 ms. See Speech-to-Speech.
  • Lower silence_timeout_ms from 800 to around 300 if your callers speak cleanly — it comes straight off the pause before every reply.
  • Point llm.provider at ollama to run the model locally, so nothing but audio transcription leaves your machine.
Prefer describing an agent to writing its config? The Agent Builder does that.

Answer a real phone call

Everything above needed no Twilio. To have the agent pick up an actual phone, you now also need a Twilio account, a number (~$1/mo) and a tunnel.
1

Expose via ngrok

2

Add the Twilio values to .env

.env
3

Bind the number

Or re-run the seed script with the number set, binding it to the agent you already talked to:
4

Call your number

The receptionist answers, understands your intent, and routes accordingly.

What You Just Built

Twilio opens a media-stream WebSocket to /ws/media-stream; the pipeline runs until you hang up, then the call is marked completed.

Manual Setup via API

If you prefer to set things up step by step:

Create a project

Project and first-API-key creation are gated by the platform credential — send X-Platform-Key matching your PLATFORM_API_KEY:

Create an API key

Create an agent

Publish the agent

Bind a phone number

Configure Twilio webhooks

Set your Twilio number’s webhook URLs:
  • Voice URL: https://xxxx.ngrok.io/webhooks/twilio/voice/inbound (POST)
  • Status URL: https://xxxx.ngrok.io/webhooks/twilio/status (POST)

Call your number

That’s it — call the number and talk to your agent.

Environment Variables

See Providers for provider-specific configuration and Video Avatar for avatars.