Voice AI · Realtime systems

Voice & Memory.

A realtime voice assistant with persistent memory and tools for calendar, CRM, and search.

A CONVERSATION, WITH CONTEXT.WHISPERspeech to textGPT-4o + MEMORYpersistent vector storeELEVENLABStext to speechCALENDAR / CRM / SEARCH<1.5sLATENCYARCHITECTURE ILLUSTRATION
DOCUMENTED OUTCOME<1.5sLatency via async streaming
TECHNOLOGIES

Python · Whisper · GPT-4o · ElevenLabs · LangChain · Google Cloud Run

The problem

A useful voice agent needs to respond quickly, retain context, and take action through external tools within the same interaction.

The engineering

01

Connected Whisper speech-to-text, GPT-4o, and ElevenLabs text-to-speech in a realtime voice agent.

02

Added persistent vector-store memory and multi-tool orchestration across calendar, CRM, and search APIs.

03

Used asynchronous streaming to reduce latency, then deployed on Google Cloud Run with CI/CD and automated quality regression tests.

The outcome

Achieved sub-1.5-second latency through asynchronous streaming, with persistent memory and tool orchestration in the deployed agent.

PROJECT SUMMARY BASED ON MY RÉSUMÉ. DIAGRAMS ILLUSTRATE THE ARCHITECTURE; THEY ARE NOT PRODUCT SCREENSHOTS.
NEXT PROJECT / 01Social Brain