Senior Software Engineer, Voice AIActive$156K

The opportunity

This is a high-autonomy, high-agency position for a voice AI engineer who thrives at the intersection of real-time systems, conversational AI, and healthcare. You'll own the architecture and delivery of Natera's Voice AI platform — a production system handling thousands of…

What you'll do

  • Own the end-to-end voice AI architecture: from Twilio media streams through LLM orchestration to TTS output and call disposition

  • Design and implement multi-agent systems using tool calling, agent handoffs,: and shared conversation state for complex patient workflows

  • Build and optimize real-time audio pipelines: WebSocket streaming, codec handling (mulaw/PCM), VAD configuration, and interruption management

  • Architect analytics and observability infrastructure for voice-specific: metrics: per-segment latency (STT/LLM/TTS), call efficacy, disposition accuracy, and ASR error rates

  • Solve voice-specific challenges: turn-taking timing, silence detection thresholds, barge-in recovery, medical term recognition, and end-to-end latency optimization

  • Integrate voice agents with internal services via secure authenticated APIs

What they're looking for

  • + years of software engineering experience, with at least 2 years building: production voice AI or conversational AI systems
  • Deep experience with voice AI pipelines: you understand the end-to-end flow from telephony through STT, LLM processing, TTS, and back to the caller, and you've solved real problems at each stage
  • Production experience with agentic architectures: multi-agent orchestration, tool calling, agent handoffs, memory/state management, and LLM-driven decision making in real-time conversation contexts
  • Strong understanding of voice-specific challenges: VAD tuning, turn-taking, interruption/barge-in handling, latency budgets, audio codec management, and the differences between voice and text-based AI UX
  • Hands-on experience with telephony systems: Twilio (media streams, SIP, IVR), or equivalent platforms with WebSocket-based audio streaming
  • Proficiency in TypeScript/Node.js with strong async programming patterns;: experience with NestJS or similar frameworks
  • Experience with STT/TTS providers (Deepgram, OpenAI, ElevenLabs, Azure: Speech) and understanding of ASR accuracy challenges (domain-specific vocabulary, noise handling)
  • Production experience with LLM APIs: OpenAI (especially Realtime API), Anthropic Claude, or equivalent; prompt engineering for conversational agents