Added
openai_live agents with tools work out of the box. An agent with tools and
no s2s.delegation block now delegates to its own llm ("mode": "client"),
with your keys, running your webhook and MCP tools — end_call, do_not_call
and transfer_call included. See
Delegation.llm.failover reaches an openai_live agent’s own-llm delegate. If the
delegate’s provider fails, delegations move to the next backup, recorded as
provider.failover like on cascade. See Delegation.openai_live agents configure their delegate with s2s.delegation.
gpt-live-1 calls no tools itself — it hands them to a delegate. The new block
names an OpenAI-hosted Responses model with its reasoning effort, service tier,
own instructions and a timeout (default 30 seconds), and gpt-live-1 is told to
delegate from your tools’ own descriptions, so no prompt rewrite is needed.
s2s.extra.backend_model keeps working. handoff_to_agent on openai_live is
now refused. See Delegation.Project credentials — each project can run on its own provider accounts.
PUT /v1/credentials/{provider} stores one key per provider account (one
Deepgram key for transcription and voice, one OpenAI key for every role),
encrypted and never returned. Every call, chat turn, post-call analysis,
knowledge-base upload and eval run uses the agent’s own key, then the project’s,
then the environment’s — so the environment becomes a fallback rather than a
requirement. GET /v1/providers lists providers per role and whether each has a
platform key and a project key. New setting: CREDENTIAL_ENCRYPTION_KEY. See
API keys.A phone number can bring its own Twilio account. Bind it with
twilio_account_sid + twilio_auth_token and TurnCall looks the number up on
that account, stores the token encrypted, and uses that account for everything
the number does. See Phone Numbers.Ambience: a room the agent sounds like it is in. Upload a sound once
(POST /v1/ambience-sounds) and any agent in the project names it:
"ambience": {"sound": "<id>", "volume": 0.3}. It plays quietly under
everything the agent says, on Twilio, WebRTC and WhatsApp voice, S2S included.
Uploads are normalized for every transport, and a sound too low to survive a
phone line is flagged at upload. See Ambience.Changed
A call that fails says so, and why. A provider that stops working mid-call (a refused key, an unknown model) used to leave the call reading ascustomer_ended_call. It now ends failed with an ended_reason of
pipeline_error, and an error.raised event says what broke. See
When a call fails.A missing provider key is caught before the call starts. If no level holds
a key for a provider the agent uses, the call is refused up front, naming it
(missing_keys), instead of failing partway into the greeting. WebRTC answers
500 pipeline_start_failed. See API keys.AWS_AGENT_CREDENTIALS_ENABLED governs agent config only. Static AWS keys
are still refused inside an agent’s aws block when it is off, but an aws
project credential is accepted — it is encrypted and outside agent config.Ambience is in the call recording. The agent’s channel carries the room, as
the caller heard it. The docs said it was left out, but it has always been in
the recording. See Ambience.Fixed
Anopenai_live agent checks in on a quiet caller. The idle check-in never
reached gpt-live-1, so a caller who went quiet heard silence until the call’s
maximum duration. gpt-live-1 now asks whether they are still there, and a second
silence ends the call as customer_silent. A silence while the agent’s delegate
is still working does not count.Ambience turned off the idle check-in on openai_live. With ambience on, the
call never registered the agent as done speaking, so a quiet caller was never
asked whether they were still there. Video avatars under ambience had the same
problem.An openai_live agent opens the call. gpt-live-1 waits for speech, so an
agent with no first_message said nothing until the caller did, and one with a
first_message could stay silent too. The session now starts with an explicit
request to speak first: the first_message, or a greeting when there is none.
gpt-live-1 delivers it in its own words, usually verbatim.An eval that failed because a provider refused the agent now says so. The
iteration’s failures lead with the provider’s error (for example Bedrock’s
“Model use case details have not been submitted”) instead of only “no response
text yet”.A call that could not start went silent. WebRTC answered the offer anyway
and connected the browser to nothing, and no call.ended was sent. It now
returns an error and sends call.ended without waiting for a recording that
cannot exist.An eval run whose project keys could not be read used the platform’s keys.
It is now errored, naming why.An OpenAI agent’s llm.api_key was ignored. It now wins over the platform
key, as it already did for Anthropic, OpenRouter and custom endpoints.Deleting a document or a knowledge base now deletes its files. Before, the
rows went and the uploaded files stayed in storage forever. Re-uploading a
filename to one knowledge base also no longer overwrites the first document’s
file. Deleting a project now also removes recordings stored through Twilio’s
recording callback.2.0.0
Evals — automated behavioural testing for agents — plus tool calling on text, per-agent pronunciations, and Pipecat 1.12. Major because two things break: six endpoints are removed and one agent config is now refused at create.⚠️ Upgrade notes
Run your migrations.alembic upgrade head. Six of them: the eval tables and
their later columns, tool_invocations.session_id, and the drop of the dead
test_suites / test_runs stub. That drop is one-way by design — its
downgrade raises rather than recreating tables whose rows described runs that
never happened./v1/test-suites and /v1/test-runs are gone. They accepted rows and did
nothing; evals replace them.Name a model on ollama, custom_openai, bedrock and openrouter. An
agent that names none is now a 422 at create. Such an agent never worked — the
model is the deployment on those providers, so it failed on its first LLM turn —
but a create that returned 201 returns 422 now, so anything provisioning
agents should set llm.model explicitly before upgrading.mcp 1.x is no longer supported — Pipecat 1.12’s mcp extra requires
>=2.1.1.Existing agents gain an idle hangup. user_idle_timeout_ms defaults to
10000, so a caller who goes quiet for ten seconds after the agent stops speaking
hears idle_message, and a second silence ends the call with
ended_reason: customer_silent. Set it to 0 to keep the old behaviour.Evals need a second process. turncall-eval-worker — same image, own
entrypoint — plus Redis; without it a run is accepted and never executed. The
default judge and persona are a local Ollama, so a simulation or an eval:
assertion needs one reachable even when the agent’s own LLM is hosted. Or set
EVAL_JUDGE_PROVIDER / EVAL_SIMULATOR_PROVIDER.Added
Evals: automated behavioural testing for agents. State what an agent must do and have TurnCall check it on every change. Scenarios are scripted (a fixed conversation with per-turn expectations) or simulations (a persona and a goal, improvised by an LLM and judged on the outcome), run in text or audio modality against the agent’s real pipeline. A dedicated worker executes them, never the API process. See Evals.Tool mocking, failing closed. A scenario carriestool_mocks and a
tool_policy; under the default mock_only a tool call with no mock is refused
and the run is errored naming it, so pointing a scenario at a real agent cannot
book a real appointment ten times.turncall eval run exits non-zero when your agent regressed, so a pull
request can be gated on behaviour. A failure is 1, an error 2, a cancellation
3, a batch still running when the command stopped waiting 4, a usage mistake
64 — distinct because “your agent regressed” and “we could not check” want
different alerts, and a judge outage must never read as a regression.eval.run.started / eval.run.completed webhooks. The completed event is
comprehensive: status, counts, every iteration’s transcript and failures, and
the configuration snapshots. The envelope gains a nullable eval_run_id beside
call_id and session_id; existing subscribers are unaffected.Choose the model that judges an eval. A scenario names
judge: {provider, model, temperature} and simulator: {...}; provider is a
closed set — ollama, openai, anthropic, jev — mapping to factories
TurnCall ships, because pipecat’s own escape hatch is a dotted import path and that is
remote code execution from a request body. EVAL_JUDGE_PROVIDER / _MODEL /
_TEMPERATURE (and the EVAL_SIMULATOR_* trio) set one for the whole platform;
the scenario’s block wins, taken whole. A judge on a run request is
refused, because a run is what it was queued as. Sending "judge": null on an
update returns a scenario to the default. See
Choosing the judge.jev as a judge — TypeSafe’s classifier, a few hundred milliseconds per
verdict with a calibrated probability rather than a model’s opinion of its own
certainty. Judge-only, and structurally so: the provider allowlist is split by
role, because a simulator is linked into a live pipeline to speak and a
classifier has nothing to link, so simulator: {"provider": "jev"} is refused
naming the role. Needs TYPESAFE_API_KEY and the evals-jev extra — neither
checked at scenario create, because both belong to a deployment and a stored
scenario outlives the one it was written against, so the run fails naming
which is missing.TurnCall decides who writes the reason. Since pipecat 1.12 the verdict comes
from a classifier and a second model — the explainer — writes the prose behind
every no. For a classifier-only judge pipecat falls back to its own default LLM
service, a local Ollama, so a judge pointed at a hosted vendor would silently
acquire a dependency on a service nobody provisioned and report it as an
unreachable connection mid-run. The explainer is compiled from the platform’s
configured judge instead, and where the platform configured none, reasons are
turned off explicitly. LLM-backed judges are left exactly alone.
harness_config records explainer_used, explainer_provider and
explainer_model.A confidence beside every verdict a judge produced — a scripted expectation
and its failures, a simulation’s goal verdict, and each per-turn verdict of a
judged metric. null is not zero: zero means the judge was certain of the
opposite. There is deliberately no aggregate. Pipecat 1.12 carries a confidence
and then drops it building the result objects TurnCall reads, so the field is
null on runs today and starts reporting on the bump that threads it through.A run carries its own warnings instead of the worker’s log — a mock keyed to
a tool the run can never call, an agent whose real tools are allowed to fire, a
scenario in which nothing can fail, a judge that changed since the last verdict,
a worker older than the code it runs. They ride on the run, the batch summary,
eval.run.completed and turncall eval show, because the person reading a green
verdict is rarely the one who wrote the scenario.A light read for a run, and paged eval lists. GET /v1/eval-runs/{id}?view=summary
answers “is it done yet” without every iteration’s transcript and all three
snapshots, and both eval lists page in the same paginated envelope the rest of
the API uses. eval.run.completed stays comprehensive: that argument is about a
subscriber who gets one event, not an endpoint on a timer.tts.pronunciations — a flat word → IPA map, the outbound mirror of
stt.keyterms. One spelling, each provider’s own markup generated from it:
Cartesia inline phonemes, ElevenLabs SSML <phoneme>, Deepgram’s inline object;
OpenAI has no pronunciation surface. Validation is TurnCall’s own, because
pipecat’s IPA helpers normalize and tokenize but never reject — so met-FOR-min
would be accepted, inserted into the markup and change nothing audible, and is a
422 naming the word. On a voice that cannot read the markup the markup is
withheld: a <phoneme> tag reaching a service that is not parsing SSML is
read out loud, measured at 3.5–4.3× the audio, which is worse than the
mispronunciation. See Pronunciations.warnings on an agent read. Two things an author needs told that cannot be
refusals: a “pronunciation” equal to its own spelling ({"Metformin": "metformin"}
— every character is legal IPA, so the validator cannot see it, and it reaches
the voice as real markup), and a voice that will not read the markup, which was
known at pipeline build and said only to a server log, mid-call, about a config
written days earlier. Derived on every read and never stored, so an agent whose
model moved out from under its pronunciations starts saying so without being
edited.A verdict says which judge decided it. harness_config gains
judge_provider and judge_temperature beside the model, service and factory
it already recorded — and a run whose judge differs from the same scenario’s
last verdict carries a judge_changed warning naming both, so a scenario that
turns red can be told apart from a judge that moved underneath it.A run records the code that executed it — worker_version,
worker_started_at, and worker_stale when the worker’s source is newer than
the process running it. A worker left running while the code moves under it
produces results that look ordinary and silently lack whatever shipped since.Derive a scenario from a real conversation — POST /v1/eval-scenarios/from-call turns a finished call into a draft, and
POST /v1/eval-scenarios/from-session does the same for an SMS, chat or
WhatsApp conversation. Each tool mock is seeded with what that tool actually
returned.POST /v1/webrtc/connect returns the call id beside sdp and type. A
browser that had just held a conversation could not say which call it was, so it
could not fetch the transcript, link the recording, or convert it. Additive — a
client reading only sdp/type sees no change.Tool calling on SMS, chat and WhatsApp text. Text conversations could not
call tools at all — an agent that books meetings on a phone call would claim it
had booked one over SMS. Works on every provider family: OpenAI-compatible,
Anthropic and Bedrock Converse each have their own dialect. Calls within a round
run concurrently, and the loop re-asks with tools withheld after five rounds so a
reply always goes out. Built-ins stay voice-only.Text tool calls are recorded. GET /v1/chat/sessions/{id}/tool-invocations.
output_json and latency_ms are populated on both voice and text, and a tool
returning {"error": ...} is recorded as failed rather than succeeded.execution_mode: "async" lets a slow tool outlive an interruption instead of
being cancelled by it. Voice only.max_call_duration_seconds now caps a call, with its own ended_reason of
max_duration_reached. interruption_enabled: false turns off barge-in on
cascade. silence_timeout_ms sets the VAD stop window. All three were
documented and did nothing.New limits you can set: MCP_MAX_TOOLS_TOTAL, MCP_CONNECT_TIMEOUT_SECONDS
and TOOL_MAX_RESPONSE_BYTES.stt.keyterms — one list of vocabulary hints (product names, SKUs,
surnames), mapped to whatever the chosen provider calls it: keyterm on
Cartesia, keywords on OpenAI, keyterms on ElevenLabs, and on Deepgram
whichever spelling the model takes. Deepgram 400s on the wrong one rather than
ignoring it, so hand-writing extra: {"keyterm": ...} on a Nova-2 agent killed
the call at connect; keyterms picks correctly and takes precedence over
extra.user_idle_timeout_ms (default 10000, 0 disables) ends a call the
caller has walked away from. After the agent stops speaking, that much silence
prompts idle_message (“Are you still there?”); a second consecutive silence
hangs up with a new ended_reason of customer_silent. Speaking resets the
count. Cascade speaks the line through TTS; S2S asks the model to check in and
words it itself.Changed
Pipecat 1.12, up from 1.9 via 1.10 and 1.11, and theopenai 3 SDK. See
adr/0019. Voicemail detection is now one processor over an explicit
LLMClassifier, nltk is replaced by sentencex, and the ollama judge is an
LLMClassifier with a 60s budget: 1.12 judges a simulation one bot turn per call
and fires those calls together against a 10s default, and a local Ollama
serializes them — measured at 18s each for four concurrent against 4.5–5.8s
alone, so every turn timed out and the run reported judge call failed, which
from outside is indistinguishable from an agent regression.An empty user turn is handled rather than dropped, and on by default. A turn
VAD opened and closed with no transcript — a cough, a door, an unrecognized word
— gets a single LLM run if it interrupted the agent, so the agent asks the
caller to repeat instead of stopping dead mid-sentence. One arriving while the
agent was already waiting is left alone: replying to a passing truck is
intrusive, and the idle timeout already covers a caller who has genuinely gone
quiet.stt.model and llm.model now mean “the provider’s own default” when
empty, instead of carrying one provider’s model to all of them. Unset,
stt.model resolves to nova-3-general / gpt-transcribe / scribe_v1 /
ink-whisper, and llm.model to gpt-4o-mini on OpenAI or claude-sonnet-5
on Anthropic. An explicit value is still passed through untouched, so a wrong
one is reported by the provider rather than silently replaced.Claude is sent no temperature, on the direct Anthropic path and through
Bedrock — current models reject it with a 400 that ends the call. Set one
deliberately on an older model with llm.extra.MCP tool-name precedence follows your configured server order rather than
whichever server answered the handshake first.Fixed
The caller sat in dead air, 2.4–5.7s per turn. VAD and Smart Turn were charged one after the other rather than together — Smart Turn’s silence counter cannot start until VAD has waited out its own window, which at the 800ms default put end-of-turn 1.8s after the caller stopped talking before the LLM was even called. With Smart Turn on, the model decides the turn and VAD uses pipecat’s 0.2s.Every eval run scorederrored — no eval feature had ever reported a verdict.
Pipecat’s harness is an RTVI client and the server half is a processor in the
bot’s pipeline plus an observer on the task; TurnCall built neither, so every
iteration died at the bot-ready handshake. Because errored counts toward no
rate, the outage read as “no signal”.A handoff could not survive the caller talking over it. handoff_to_agent
queues the target’s prompt and tool schema as frames, and pipecat drains
interruptible frames on barge-in — so a caller saying “okay, thanks” mid-handoff
left the model running the previous agent’s prompt, advertising the previous
agent’s tools, over a context already wiped. Nothing raised and nothing was
logged. The switch is also atomic against cancellation now: an interruption
cancels in-flight sync tool calls, and a DB round-trip after the commit left the
call record handed off while the live pipeline kept the previous agent.tts.pronunciations could never work on ElevenLabs. Pipecat reads SSML
<phoneme> tags only with a phoneme-capable model and enable_ssml_parsing, a
constructor argument no tts.extra could reach. TurnCall passes it for exactly
the agents that have pronunciations and such a model — never for everyone, since
SSML parsing changes how all of that agent’s text is read.Cancelling a running eval run now stops it, and stays cancelled. DELETE set
the column and nothing else: every remaining iteration ran and was paid for, and
the finisher then wrote passed over cancelled and announced
eval.run.completed for a conversation the user had asked to stop.A hung eval iteration is bounded, and the janitor’s cutoff is derived from the
run’s iterations rather than a flat 900s. Both were the same false alarm from
opposite ends: an unbounded harness call held a worker slot for the life of the
process, and a flat sweep reclaimed healthy multi-iteration runs mid-flight —
which the CLI reported as a regression.A tag’s fan-out is capped. One POST with a popular tag and iterations: 50
queued scenarios × iterations full conversations, so a project with 200
scenarios on pre-publish was 10,000 conversations away from one HTTP call.
EVAL_MAX_SCENARIOS_PER_REQUEST (50) refuses it naming the count, never
truncating.The CLI’s deadline fits the batch, and a batch that outran it exits 4, not
2. The default was a per-run budget used as a whole-batch one, so ten scenarios
behind a four-slot worker made the command give up on a healthy suite and report
an outage. “We stopped waiting” and “nobody could tell” are now different codes,
because they have different remedies.A scenario that cannot fail is warned about. Validation checks shape, not
strength, so a scenario whose every expectation is a content-free event parses,
stores, runs and reports passed forever against an agent whose LLM returns
nothing — after a provider 404 pipecat still emits an empty llm_response. The
create/update response and the run both carry scenario_cannot_fail, and such a
run scores failed, never errored.The eval worker’s readiness probe stopped looking like a failure. Every
iteration logged opening handshake failed with a traceback and nothing was
wrong. The cost was not the noise: a probe succeeding and a bot that could not be
reached produced the same traceback.A sentence the STT split in two became two scenario turns in from-call, and
audio evals work in Docker — pipecat’s TTS cache default lives under $HOME,
which a container loses on recreate, so EVAL_TTS_CACHE_DIR exists and is
mounted.MCP servers now connect on WebRTC and WhatsApp voice, and S2S advertises the
tools it was already holding sessions for. A tool name claimed by two servers no
longer makes the request invalid, and the tool that runs is the one the model was
shown. handoff_to_agent now swaps tools with the prompt.An inline agent from call-init crashed Twilio calls, and on every transport
skipped post-call processing — which is what dispatches call.ended, so such a
call completed and told nobody.Tool results are capped before they enter the context, for webhook tools as
well as MCP.A default google S2S agent sent OpenAI’s model to Gemini. s2s.model and
s2s.voice default to OpenAI Realtime’s values whatever the provider, so an
agent naming neither reached Gemini holding gpt-realtime-2.1 and alloy.
Gemini closed the socket with a 1008 and the call died at connect. Both
sentinels are now swapped for Gemini’s own (gemini-3.8-live, Charon), the
way the aws and openai_live paths already did.The model reported in call.ended is the one that actually ran — post-call
analysis resolved its own default separately from the pipeline, so the two could
disagree.Security
An agent’s secrets leaked through an eval run. The config masker guarded one exit — the agent response builder — and the eval feature added two more paths carrying the same config:GET /v1/eval-runs/{id}, where a viewer key read every
credential in the clear, and eval.run.completed, which POSTed them to a webhook
subscriber. An inline target leaked the same way. The masker now lives where no
exit can miss it, and a run’s stored resolved_config is masked on the way in.A judge or simulator cannot name a factory. Pipecat’s escape hatch is a
dotted path handed to importlib.import_module — fine for a file on disk, remote
code execution from a request body. One naming a factory anywhere pipecat reads it
is refused at the API boundary, and api_key is not a field.The BYOM/MCP URL allowlist could be talked past. * in an fnmatch pattern
spans /, so https://*.trusted.com/* also accepted
https://evil.com/x.trusted.com/y. A pattern naming a host must now match the
URL’s host on its own. MCP urls go through the gate too — they had none.1.1.0
Bedrock and Nova Sonic, GPT-Live-1, Pipecat 1.9, and three config fields that were accepted and then ignored. No REST API, agent config schema or webhook payload shape changed.AWS Bedrock and Amazon Nova Sonic
New LLM provider:bedrock — run Anthropic, Meta, Mistral or Amazon foundation models through AWS, on voice and on SMS/chat. Model ids pass through verbatim, so direct ids, cross-region inference profiles (us.-prefixed) and provisioned-throughput ARNs all work. llm.extra forwards to Bedrock’s additionalModelRequestFields, which is how Anthropic extended thinking is reached.New S2S provider: aws — Amazon Nova Sonic 2 (amazon.nova-2-sonic-v1:0), with voice defaulting to matthew and s2s.extra.endpointing_sensitivity (LOW/MEDIUM/HIGH) tuning how quickly it decides the caller stopped speaking.New aws agent config block — one per agent, shared by both providers, so an agent’s LLM and voice leg always authenticate as the same principal:role_arn, static keys, a named profile, or the ambient chain (env vars, SSO cache, EC2 instance profile, ECS task role, EKS IRSA). Nova Sonic sessions roll over roughly every 6 minutes and credentials are re-resolved each time, so calls outliving a temporary credential keep working.New setting: AWS_AGENT_CREDENTIALS_ENABLED (default false) — per-agent static AWS keys persist in the agent config, which is not encrypted at rest, so agents supplying them are rejected at create unless this is enabled. role_arn needs no flag and stores no durable secret. AWS_REGION now also sets the default Bedrock region; set aws.region per agent, since model availability rarely matches where your S3 bucket lives.See Providers, Speech-to-Speech, examples/bedrock/, and adr/0016.The system prompt moved onto the model, not the conversation
Internal change, no config or API difference. An agent’ssystem_prompt — with guardrails and any prompt-mode knowledge preamble — is now set on the LLM service itself rather than inserted as the first message of the conversation. Pipecat deprecated the message form in 1.9 and stops honouring it in 2.0.Two things follow. The conversation context now begins empty, so a transcript no longer carries the instructions as a phantom first turn. And handoff_to_agent switches agents by updating the model’s instruction instead of rewriting that message — which had to change at the same time, because the instruction is prepended to the context, so doing both would have sent the previous agent’s prompt alongside the new one.STT: provider passthrough
stt.extra now reaches the provider. The field has existed since the beginning and was silently ignored — the API accepted it, stored it, and dropped it. Keys naming a setting TurnCall already manages (model, language, and Deepgram’s punctuate / smart_format / interim_results) are ignored rather than applied, so extra cannot quietly override your configuration.On Deepgram it unlocks the rest of the transcription surface: profanity_filter, diarize, redact, keyterm, version, utterance_end_ms.Voice pipeline upgraded to Pipecat 1.9
Upgraded from 1.8.1. No config or agent changes required.Provider defaults refreshed — these apply only to agents that do not set the value explicitly. Pin it in your agent config to keep the previous behaviour:OpenAI shuts down
gpt-4o-transcribe on 2027-02-26 and gpt-transcribe replaces it. LiveAvatar has deprecated VP8.Deepgram transcripts are no longer profanity-filtered. Deepgram’s filter rewrites the words it matches rather than tagging them, so one false positive silently altered a transcript you store, analyse, and receive in the call.ended payload. Transcripts now carry what the caller actually said. Set stt.extra.profanity_filter: true on an agent to turn the filter back on.OpenAI STT no longer clips the last word. Segment-based STT now hears 0.5s of trailing silence before transcribing, so the final word of an utterance is transcribed instead of dropped or garbled. The padding is sent to the provider, so it counts toward your STT usage.Assistant transcripts no longer contain text the model never wrote. A word-timing event that matched nothing left to speak was previously added to the conversation as spoken text. Nothing is lost — the text it should have covered is carried by the word that puts the sentence back in step.Fixed: ending a call during voicemail detection. While the voicemail classifier’s gate was closed, the frame that ends a call could be dropped instead of reaching the pipeline, leaving the caller on the line until the idle timeout. An agent calling end_call in that window now ends the call.TTS: speed, provider passthrough, and a newer Cartesia default
tts.speed now works on every provider. It was read only for Cartesia, so setting it on the default Deepgram voice did nothing. Deepgram (0.7–1.5), OpenAI and ElevenLabs all honour it now. Leaving it at 1.0 sends nothing, exactly as before.tts.extra now reaches the provider. Previously only Cartesia’s language and emotion were read and every other key was dropped. Keys naming a setting TurnCall already manages — voice, model, speed — are ignored rather than applied, so extra cannot quietly override your configured voice; set tts.voice for that.Cartesia’s default model is now sonic-3.6, matching Cartesia’s current Sonic release. It applies only to agents that leave tts.model empty; pin "sonic-3.5" to stay on the previous one.Fixed: a Deepgram voice was being sent to other providers. tts.model and tts.voice both default to aura-2-helena-en — a Deepgram Aura voice — whatever provider you set. An agent that chose Cartesia, OpenAI or ElevenLabs without also naming a model and voice sent that Aura name to the wrong provider. Each provider now falls back to its own default: tts-1/alloy for OpenAI, eleven_flash_v2_5/Rachel for ElevenLabs, sonic-3.6 for Cartesia. Cartesia voice ids are account-specific, so an unset Cartesia voice now logs a warning naming the setting rather than sending something wrong.Knowledge retrieval is guidance, not a system prompt
Chunks retrieved inauto mode are now attached as developer rather than a second system message. They are guidance for the turn, not the agent’s instructions, and a system message mid-conversation was already being downgraded to user by every provider except OpenAI — so this is no worse anywhere and correct on OpenAI. Your agent’s own system_prompt is unaffected.New S2S provider: OpenAI GPT-Live-1
s2s.provider: "openai_live" runs OpenAI’s gpt-live-1, announced on 10 September 2026. It is full duplex — it listens and speaks at once, and decides for itself when to answer and when to stop on being talked over — where the Realtime API is turn-based. OpenAI reports a 30-point gain on Full Duplex Bench over gpt-realtime-2.1, and calls out telephony as a target.s2s.extra.backend_model and OpenAI hosts it; the function calls it makes run against your agent’s own tools. Left unset, the live model answers alone.Two differences from openai: s2s.temperature is supported here (Realtime GA rejects it), and s2s.turn_detection must stay server_vad — client-side VAD would fight a model doing its own turn-taking, so pipecat_vad is rejected at config time.1.0.0 — first open-source release
TurnCall is now open source under the MIT licence, with container images published toghcr.io/kobikis/turncall. From this release on, a breaking change to the REST API, the agent config schema, or the webhook payloads means a major version bump.Changed
Voice pipeline upgraded to Pipecat 1.8.1 — from 1.5. No config or agent changes required.Provider model defaults refreshed — these apply only to agents that do not set a model explicitly. Pin the old value in your agent config to keep it:A call now ends when a service can no longer work — Pipecat 1.8 stops using an STT, TTS or LLM after an unrecoverable failure (a rejected API key, an unknown model or voice, a connection that will not re-establish) rather than retrying per chunk. TurnCall ends the pipeline promptly instead of leaving the caller on a silent line until the idle timeout.
Bug fixes
Per-agent Deepgram config was ignored —stt.model and stt.language were hardcoded to nova-2 and en for the Deepgram provider, discarding whatever an agent specified. Both are now honored. If you run non-English agents on Deepgram, their configured language now actually takes effect — verify transcription quality after upgrading.New features
Temperature and max tokens on voice calls —llm.temperature and llm.max_tokens were accepted in agent config (and honored on chat/SMS) but silently ignored by the voice pipeline; they now apply to every cascade voice call across all LLM providers (OpenAI, Anthropic, Ollama, custom endpoints, OpenRouter). S2S agents gain the same knobs on the s2s block: s2s.max_tokens on both providers, s2s.temperature on Gemini Live (the OpenAI Realtime GA API has no temperature control — setting it there returns a clear validation error instead of being silently dropped). The voicemail classifier stays pinned to a deterministic low temperature regardless of agent settings.Platform credential for bootstrap — Project creation and first-API-key creation (POST /v1/projects, POST /v1/api-keys) are now gated behind a single privileged credential: the X-Platform-Key header must match the server’s PLATFORM_API_KEY. The gate fails closed — unset means every bootstrap call is rejected — so anonymous callers can no longer mint projects on an exposed deployment. Everything after bootstrap keeps using project-scoped tc_... keys. env.example ships a dev default (dev-platform-key); set a strong unique value in production.One-command example launchers — Every example now ships a run.sh: it reads the shared values (TURNCALL_NUMBER, TWILIO_PN_SID, PUBLIC_BASE_URL) from .env, names exactly what’s missing if unset, and passes extra flags through to the example’s setup.py. The seed script and all examples also send the new platform credential automatically.Prohibited topics enforced — guardrails.prohibited_topics is now compiled into the system prompt with refusal instructions, so listed topics are actually declined instead of being config-only metadata.Knowledge injection for voice prompt mode — Knowledge bases attached in prompt mode now inject their document text into the system prompt on voice calls too (previously chat-only), and agents with any knowledge attachment get a hint that they have a knowledge base, reducing “I don’t have access to that” refusals.Bug fixes
S2S calls now produce transcripts — Speech-to-speech pipelines (OpenAI Realtime, Gemini Live) were not persisting conversation transcripts; they are now captured to the call record andtranscript.final events like cascade calls.WebRTC calls finalized as FAILED on pipeline build errors — A WebRTC call whose pipeline failed to build was left dangling in in_progress; it now finalizes as FAILED with a proper call.ended.Validation errors return 422, not 500 — Request bodies that failed model_validator checks surfaced as generic 500s; they now return 422 with the validation detail.New features
Speech-to-Speech via gateways (Grok, and more) — Theopenai S2S provider now takes an optional s2s.base_url pointing at any OpenAI-Realtime-compatible gateway (Vercel AI Gateway, LiteLLM) or xAI direct. Provider-prefixed realtime models like xai/grok-voice-think-fast-1.0 and openai/gpt-realtime-2 then stream over the same WebSocket protocol — no new provider, same low-latency pipeline. The endpoint is SSRF-gated by BYOM_ALLOWED_URL_PATTERNS, and a gateway base_url lifts the OpenAI voice allowlist so third-party voices pass through. See the Speech-to-Speech guide.Updates
Voice pipeline upgraded to Pipecat 1.5 — Under-the-hood upgrade of the real-time voice engine, bringing upstream latency and resource-teardown fixes. Fully backward compatible — no config or agent changes required.Bug fixes
Gemini Live default model refreshed — The Speech-to-Speech example and docs now default thegoogle provider to models/gemini-3.1-flash-live-preview; the previous gemini-2.5-flash-native-audio-preview is on Google’s deprecation schedule. Gemini voices are no longer allowlisted — its native-audio voice set grows per model, so any voice is accepted and Gemini validates it on connect.Agents no longer read markdown aloud — Language models sometimes format replies with markdown (**bold**, `code`, # headings), which text-to-speech would voice literally (“asterisk asterisk”). Markdown symbols are now stripped before speech across every TTS provider (Deepgram, ElevenLabs, OpenAI, Cartesia), so agents speak the words, not the punctuation.New features
Takeaways — reusable structured outputs — Define a named JSON schema once (optionally with a custom prompt and model), attach it to any agents viaanalysis.takeaway_ids, and every call ends with validated JSON extracted from the conversation — CSAT scores, lead fields, booking details. Each takeaway runs as its own concurrent extraction (one failure never affects the others), results are schema-validated with automatic retry, and land keyed by name in call.ended under analysis.takeaways and in the analysis API. CRUD at /v1/takeaways. See the takeaways guide.Hybrid knowledge retrieval — Knowledge base search now combines vector similarity with Postgres full-text search, fused by rank (RRF). Exact-term questions against record-like documents (“what is the flight date?”, reservation codes) hit reliably even where embeddings under-score; retrieval degrades to lexical-only if embedding generation fails. One migration adds the index.Contextual chunk enrichment — At ingest, each chunk is prefixed with an LLM-written sentence situating it in its document (filename, what this part covers, key entities), improving retrieval on multi-document knowledge bases. Best-effort: failures fall back to a filename prefix and never block an upload.GET /v1/calls/{id}/recording — Fetch a call’s recorded WAV directly from the API. The endpoint streams the audio from object storage, so you can download or embed a recording without minting a signed URL yourself.Updates
PDF extraction cleanup — Web pages saved as PDF no longer drown retrieval in navigation links, session URLs, and repeated page headers; the noise is stripped at extraction.Smarter voice retrieval queries — Auto-mode retrieval now windows the query over the previous user turn and the agent’s last reply, so follow-ups like “and what time?” carry the entities they refer to.Bug fixes
duration_ms missing when the agent ended the call — Calls ended by the end_call tool could finalize with a null duration_ms. Duration is now stamped on every end path, so call.ended and the call record always carry it.Agent-spoken lines missing from transcripts — The agent’s first_message, voicemail message, and transfer messages are spoken directly (not through the LLM turn), and were absent from the stored transcript and transcript.final webhooks. All directly-spoken utterances now appear in the transcript alongside conversational turns.Transcript speaker field always null — On transcript.final events derived from finalized speech, the speaker field was populated from the wrong source and came through as null. It now correctly reports role (customer / assistant).New features
Signed tool webhooks — Custom webhook tools can now set awebhook_secret; when present, TurnCall HMAC-signs each tool POST with X-TurnCall-Signature / X-TurnCall-Timestamp using the same v1= scheme as event webhooks, so your endpoint can verify the call really came from TurnCall. Unset = unsigned, fully backward compatible.PUT /v1/phone-numbers/{id} — Update a number’s routing, server_url, or sms_enabled in place. The phone id and its call-init server_url_secret stay stable across edits — no more unbind/rebind rotating the secret your call-init endpoint verifies with.DELETE /v1/agents/{id} — Delete (archive) an agent. Call history, transcripts, and analyses remain queryable.Updates
Knowledge retrieval default threshold 0.7 → 0.3 — The defaultsimilarity_threshold for knowledge base search and agent attachments was calibrated for older embedding models. text-embedding-3-small (the default) scores related content in the 0.3–0.5 range, so the 0.7 default filtered out everything. Existing attachments keep their stored threshold — re-link (or set similarity_threshold explicitly) to pick up the new default.Bug fixes
Calls stuckin_progress after caller hangup — When a caller ended the call (hanging up on Twilio, closing the tab on WebRTC, or ending a WhatsApp session), the call could stay at status=in_progress with no ended_at, duration_ms, or post-call analysis, and the call.ended webhook never fired. Caller-initiated hangups now finalize the call on every transport and deliver call.ended exactly once, whether the call ends via the caller, the end_call tool, or a Twilio status callback.New features
OpenTelemetry tracing & pipeline observers — Every call is now instrumented for per-stage latency. Tracing emits a span tree (conversation → turn → STT/LLM/TTS, with TTFB and token counts) to any OTLP backend (Jaeger, Grafana Tempo, Datadog, …) — the trace’sconversation_id is the call_id, so a trace links straight to the call. Five built-in observers log latency, turn timing, LLM, transcription, and startup cost. Both are on by default and cover cascade and S2S. Point OTEL_EXPORTER_OTLP_ENDPOINT at a collector to see traces; tracing self-disables in production without an endpoint (it never console-exports on the audio path).Warm call transfer with operator briefing — transfer_call now does real warm transfers, not just blind ones. Set transfer_mode: "warm" and the operator hears a briefing before the caller is bridged — either a literal string or {"from_summary": true} to summarize the conversation on the fly. Both modes can play a transfer_message to the caller first (“Connecting you to support…”), and a fallback_message covers the operator not answering. Works from the agent (the transfer_call tool) and the control API (POST /v1/calls/{id}/transfer). See the tools guide and the examples/call-transfer example.Transfer answering-machine detection — when a transfer’s destination answers, a new transfer.answered webhook reports answered_by (human / machine), so you can tell when a transfer reached voicemail. (The caller is still connected and can leave a message — voicemail is detected, not blocked.)Warm transfer and the no-answer fallback require PUBLIC_BASE_URL to be set (Twilio calls back to TurnCall for the briefing and fallback). Cold transfer and the caller message work without it.New features
agent_id and event_id on every webhook — The delivered webhook envelope now carries agent_id (the agent that handled the event, resolved from the call’s current active agent so handoffs are reflected) and event_id (a unique id, stable across delivery retries and shared across subscribers — use it as a deduplication key). Both sit at the top level alongside call_id and session_id.ended_reason on call.ended — The end-of-call webhook now reports why a call ended, distinct from the coarse status: customer_ended_call, assistant_ended_call, customer_did_not_answer, customer_busy, voicemail, transferred, pipeline_error, telephony_failed, or unknown.Richer call.ended payload — call.ended now also includes status (final call status), provider_call_sid (correlate with Twilio), and metadata (the custom data you attached at call-init, echoed back for CRM correlation).Updates
tool.result includes the tool output — The tool.result webhook payload now carries the tool’s result, not just its name and arguments.Breaking changes
Transcript events userole, not user_id — On transcript.final events the speaker field was renamed from user_id to role (values customer / assistant) — it is a speaker label, never an identifier. Update consumers that read payload.user_id.Bug fixes
WebRTC calls in Docker — The runtime image now installs the native libraries (libxcb, libGL, glib) that the WebRTC media stack loads at runtime, fixing webrtc/connect failures (libxcb.so.1: cannot open shared object file) on self-hosted deployments.New features
Built-in call recordings on every transport — TurnCall now records every call itself and writes a WAV file to your configured storage, whether the call comes in over Twilio, WebRTC, or WhatsApp. No Twilio recording configuration is required. When the file lands,recording_url is populated, recording_status flips to completed, and a recording.ready event fires.Reliable call timestamps — started_at and duration_ms are now stamped by the call pipeline itself instead of relying on Twilio status callbacks. duration_ms is always computed from ended_at - started_at, so it stays accurate even when carrier callbacks are delayed or dropped.Updates
Smaller, hardened Docker image — The official Docker image is now a multi-stage build that ships only runtime dependencies, runs as a non-root user, and skips the ~2.5 GB of CUDA libraries that were previously pulled in by default. Self-hosters get a leaner image with a smaller attack surface and no changes required to deploy. See Quickstart.Pipeline metrics enabled by default — Call pipelines now emit timing and usage metrics out of the box, giving you visibility into per-stage latency and provider usage for every call.call.ended now waits for the recording — The call.ended webhook is gated on both post-call analysis and recording persistence, so the payload carries a populated recording_url on the happy path instead of a stale null. The event always fires, and now includes recording_status so subscribers can distinguish a failed recording from one that’s still uploading.Bug fixes
Empty recordings on inbound Twilio calls —recording_url is no longer blank and recording_status no longer stays stuck at none. Inbound Twilio calls use a media-stream connection that never triggered Twilio-side recording, so no file was ever produced. The pipeline now records the call directly.Scrambled audio on PSTN calls — Fixed a resampler bug that caused clicks and aliasing on continuous TTS audio over Twilio calls. Outbound audio is now clean across frame boundaries.Mid-call dead air and cut-offs — Transcript taps and webhook delivery no longer run inline on the realtime audio path, so slow webhook endpoints or database writes can no longer cause brief audio stalls or words being cut off mid-sentence.Dropped call events under load — Resolved a race condition that could cause concurrent transcript, handoff, and lifecycle events to collide on the same sequence number and be dropped. High-throughput calls now record every event in order.Twilio webhooks rejected behind a tunnel or proxy — Inbound Twilio webhooks no longer return 403 twilio_invalid_signature when TurnCall runs behind ngrok, a load balancer, or a container with forwarded headers. Signature validation now uses the public forwarded URL, matching what Twilio signs.New features
Video avatars — Render your agent as a photorealistic talking head during WebRTC calls. Choose between HeyGen and Tavus by settingavatar.provider on the agent config. Tavus delivers sub-600ms latency at 1080p; HeyGen streams alongside your existing voice pipeline. See Video Avatar for setup and field reference.OpenRouter LLM provider — Route LLM traffic through OpenRouter to access hundreds of models with automatic fallback routing. Configure primary and fallback models per agent to improve reliability when an upstream model is degraded. See providers.Interactive API reference — You can now explore and test every TurnCall endpoint directly from the docs. The new API reference includes request and response schemas, example payloads, and a built-in playground.Open source under MIT license — TurnCall is now fully open source. The entire project is available under the MIT license, so you can self-host, fork, and contribute freely.MCP server support for tools — You can now connect MCP servers to your agents for auto-discovered tool calling. Any tools exposed by your MCP server are automatically available during calls.Post-call analysis — TurnCall now automatically generates a structured post-call analysis after every call, including a summary, sentiment score, success evaluation, and custom data extraction. The call.ended webhook is enriched with the full transcript, recording URL, and analysis results.Agent versioning — Publish immutable agent versions, auto-promote phone numbers to the latest version, and roll back instantly when needed.A/B testing — Route traffic across agent versions with weighted A/B testing on phone numbers. Routing is deterministic by caller, so the same caller always reaches the same version.Cartesia STT/TTS provider — Cartesia is now available as a speech-to-text and text-to-speech provider, giving you another option for voice quality and latency tuning.Anthropic Claude as LLM provider — You can now use Anthropic Claude models as the LLM provider for your agents, alongside OpenAI and Ollama.Knowledge base with RAG — Upload documents to a knowledge base and attach it to agents. Three retrieval modes are available: prompt injection, automatic retrieval, and tool-based lookup.Pre-call init hook — Use the call-init server event to dynamically resolve agent configuration before the pipeline starts. You can also hand off mid-call between agents using the built-in handoff tool.SMS and chat support — Agents can now handle text-based conversations over SMS and the Chat API. Sessions are managed automatically so returning users pick up where they left off.WhatsApp Business integration — Connect your agents to WhatsApp for both voice calls and text messages through the WhatsApp Business platform.Speech-to-speech mode — A new speech-to-speech pipeline delivers ultra-low-latency voice interactions powered by OpenAI Realtime and Gemini Live, bypassing the traditional STT → LLM → TTS chain.Bring Your Own Model (BYOM) — Point your agents at any OpenAI-compatible endpoint to use custom or self-hosted LLMs as the provider.WebRTC support — Launch browser-based voice calls directly from your web application without requiring a phone number.Smart Turn and voicemail detection — Improved turn-taking with Smart Turn V3 and Silero VAD reduces false interruptions. Incoming calls are now automatically screened for voicemail so your agent can hang up early instead of talking to a machine.Updates
Pipecat 1.4 upgrade — The underlying voice pipeline has been upgraded from Pipecat 1.0 to 1.4 for improved stability and provider compatibility. No action required.Streaming audio and latency fixes — Audio resampling now drops empty buffers instead of pushing silence, reducing artifacts and lowering end-to-end latency on cascade pipelines.Rebrand to TurnCall — The project has been renamed from Voicey to TurnCall. All API endpoints, configuration files, and documentation now use the TurnCall name consistently. No action is required if you are using the hosted API.Richer webhook payloads — Thecall.ended webhook now includes call metadata (from/to number, direction, duration), the full transcript, and the recording URL. All call events are dispatched to webhook subscribers.Call recording storage — Twilio call recordings are now automatically downloaded and stored locally or in S3.Pipecat 1.0 migration — The underlying voice pipeline has been upgraded to Pipecat 1.0, improving stability and enabling new provider integrations.Renamed “assistant” to “agent” — All API endpoints and documentation now use “agent” consistently. The /v1/assistants endpoints have been replaced by /v1/agents.Bug fixes
Duplicatecall.started events — Fixed an issue where call.started was fired twice per call.Transcript sequencing — Transcript events now use database sequence numbers, preventing ordering collisions in high-throughput calls.Webhook payload format — Fixed the webhook event key to use event consistently (previously some payloads used event_type).