The opportunity
This is a high-autonomy, high-agency position for a voice AI engineer who thrives at the intersection of real-time systems, conversational AI, and healthcare. You'll own the architecture and delivery of Natera's Voice AI platform — a production system handling thousands of…
What you'll do
Own the end-to-end voice AI architecture: from Twilio media streams through LLM orchestration to TTS output and call disposition
Design and implement multi-agent systems using tool calling, agent handoffs,: and shared conversation state for complex patient workflows
Build and optimize real-time audio pipelines: WebSocket streaming, codec handling (mulaw/PCM), VAD configuration, and interruption management
Architect analytics and observability infrastructure for voice-specific: metrics: per-segment latency (STT/LLM/TTS), call efficacy, disposition accuracy, and ASR error rates
Solve voice-specific challenges: turn-taking timing, silence detection thresholds, barge-in recovery, medical term recognition, and end-to-end latency optimization
Integrate voice agents with internal services via secure authenticated APIs
What they're looking for
- + years of software engineering experience, with at least 2 years building: production voice AI or conversational AI systems
- Deep experience with voice AI pipelines: you understand the end-to-end flow from telephony through STT, LLM processing, TTS, and back to the caller, and you've solved real problems at each stage
- Production experience with agentic architectures: multi-agent orchestration, tool calling, agent handoffs, memory/state management, and LLM-driven decision making in real-time conversation contexts
- Strong understanding of voice-specific challenges: VAD tuning, turn-taking, interruption/barge-in handling, latency budgets, audio codec management, and the differences between voice and text-based AI UX
- Hands-on experience with telephony systems: Twilio (media streams, SIP, IVR), or equivalent platforms with WebSocket-based audio streaming
- Proficiency in TypeScript/Node.js with strong async programming patterns;: experience with NestJS or similar frameworks
- Experience with STT/TTS providers (Deepgram, OpenAI, ElevenLabs, Azure: Speech) and understanding of ASR accuracy challenges (domain-specific vocabulary, noise handling)
- Production experience with LLM APIs: OpenAI (especially Realtime API), Anthropic Claude, or equivalent; prompt engineering for conversational agents