Latency is the whole game
If the round trip from speech to response crosses ~500ms, the illusion breaks. Everything else is downstream of that number.
The stack
LiveKit for the media transport, Gemini Live for the model, and a thin agent worker that streams both directions. No polling, no turn-taking hacks.
The details that matter
Barge-in, a warm greeting that starts before the user finishes their sentence, and a voice that is premium-HD rather than the flat text-to-speech everyone else ships.