Platform architecture
The enterprise voice AI platform your call center is actually missing.
Zumu is not a voice bot wired into your phone line. It is the call layer, the agent, the knowledge, and the operations surface your team already needs, engineered as one platform instead of stitched together after the fact.
- Our own SIP-optimized telephony underneath
- A prompt cache that stays warm
- A knowledge graph, not a pile of embeddings
- A wallboard your supervisors actually run from
How the platform is built
Four layers, engineered as one system
Call layer, agent runtime, knowledge, and the operations surface, each one real enough to stand alone and wired together closely enough that a handoff never leaves the room.
One call runs through all four layers below, start to finish, without ever leaving the room.
The four layers
What each layer actually does
None of this is a marketing gloss on top of a plain phone line. Every claim below traces to a control that exists in the product today.
- 01
Call layer
The call arrives as a room, not a webhook
Inbound and outbound calls run on Zumu’s own SIP-optimized infrastructure, over a SIP trunk you already own or an API-native carrier. Your own staff sit inside the same system as real SIP extensions, with a role and a voicemail box, instead of a phone system your platform merely calls out to.
Escalation bridges a human into the same LiveKit room the caller is already in. Recording and live listen-in continue through the handoff, because it never was a second call.
- 02
Agent runtime
The model, warmed before it has to think
An agent is a stateless process that loads your prompt, voice, model, and tools fresh from the database at the start of every call. Underneath it, a warm prompt cache and a prewarmed worker are the difference between a fast agent and a slow one.
A steady 0.98 to 0.99 prompt-cache hit ratio in production keeps a large system prompt warm across calls, and English plus Spanish run live today on a multilingual stack built to carry more.
- 03
Knowledge
A graph the agent argues with itself before it answers
Documents, call audio, and your own website feed an eight-stage pipeline that includes a genuine adversarial debate step, not just a summarizer. The result writes to three stores at once, so retrieval is never only a vector search guessing at similarity.
Postgres holds the facts, Neo4j holds how they relate, and Qdrant holds what sounds similar. Hybrid retrieval blends all three with reciprocal rank fusion before an agent ever sees an answer.
- 04
Operations surface
The layer your supervisors actually work from
A live wallboard, listen-in, whisper, and barge-in sit next to the call, not in a separate tool your team has to tab over to. Every transfer gets a briefing before it happens and a grade after, on the human leg too, not only the AI's.
Supervisors can still listen in after your own person takes a transferred call, because the room they joined never closed.
Built to be integrated with, not just trusted
The controls your own engineers will ask about
Run your call center from your own tools, not only from ours.
- MCP
MCP runs in both directions
Zumu is an MCP client, so your own MCP servers can hand new tools to an agent, and a real OAuth 2.0 MCP server, so you can search transcripts and update a prompt from Claude Code or Cursor.
MCP server - Webhooks v2
Webhooks built for a production integration
Signed with HMAC and a key-rotation overlap so a rotation never drops a delivery. Retried, dead-lettered, and replayable, with flag rules an LLM grades against a plain-language description instead of a regex.
Webhooks - BYOK
Your own provider keys, not ours
Bring your own model, speech, and voice keys. Each one is stored with AES-256-GCM envelope encryption, validated with a live ping before it goes into service, and can point at your own base URL.
Model freedom
Not four vendors
1
The number of platforms it takes. Call layer, agent runtime, knowledge, and operations surface, shipped as one contract instead of four separate integrations.
- 4
- Layers, engineered together rather than stitched together after the fact
- 8
- Stages in the knowledge pipeline before anything reaches the graph
Every layer above is real enough to read about on its own page. None of them only work when the others do too.

Nobody is watching the wallboard at 2 AM. The platform still is.
Questions
What engineers ask about the architecture
Still have questions?
Bring your hardest integration question. We will answer it against the real API, not a deck.
Book a demo
See the whole platform, not a slide about it
Bring a call you wish had gone differently. We will walk it through the call layer, the runtime, the knowledge it drew on, and the operations surface that would have caught it.
- No credit card for the demo
- Every layer, in one session