Skip to content
Back to all notes
AI & Automation5 min read

Devesh JoshiCo-founder, product

Nine years building AI platforms serving 12,000+ engineers. LLM platforms, agentic systems (MCP), and multi-model safety evaluation.

Deploying sub-500ms voice AI intake with real-time audio streaming

Engineering low-latency speech-to-speech telephony agents using WebRTC streams and real-time audio models for seamless customer service intake.

Eliminating speech-to-text pipeline latency

Traditional voice bots chained three separate steps (Speech-to-Text -> LLM Prompt -> Text-to-Speech), introducing 2 to 3 seconds of awkward latency. Direct speech-to-speech streaming models achieve human-like sub-500ms conversational response times.

Topic Focus & Target Concepts

OpenAI Realtime API voice botsub 500ms voice AI latencyWebRTC speech to speech botAI phone support engineeringreal time voice AI intake

This is the work behind our AI and automation practice — agents with real grounding, voice intake, and retrieval that answers from your records rather than the model's training data.

Agents with real tool access

Running into this in your own stack? Twenty minutes, no deck.

Book the call