Staff Software Engineer- Foundation Model InferenceActive$190K
The opportunity
At Databricks, we are passionate about enabling data and AI teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and…
What you'll do
Build LLM infrastructure powering large-scale inference workloads for: customers through partner models (OpenAI, Anthropic, Gemini) and self-hosted models (Qwen, GPT-OSS, Llama)
Improve reliability, latency, and efficiency of distributed AI workloads
Collaborate with platform, infra, and ML teams to deliver seamless end-to-end experiences
Shape how developers and data scientists build and interact with AI on Databricks
+ years of experience in backend or infrastructure engineering
Experience with distributed systems, scalable APIs, or cloud-native infrastructure
What they're looking for
- Experience with real-time serving, ML infrastructure, or GPU orchestration
- Familiarity with service-oriented architecture, deployment pipelines, and system observability
- Exposure to platforms like SageMaker, Vertex AI, or Azure ML
- Contributions to OSS projects like MLflow, PyTorch, Ray, vLLM, SGLang