Staff Software Engineer - GenAI inferenceActive$191K
The opportunity
Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles.
What you'll do
Own and drive the architecture, design, and implementation of the inference: engine, and collaborate on model-serving stack optimized for large-scale LLMs inference
Partner closely with researchers to bring new model architectures or features: (sparsity, activation compression, mixture-of-experts) into the engine
Lead the end-to-end optimization for latency, throughput, memory efficiency,: and hardware utilization across GPUs, and accelerators
Define and guide standards to build and maintain instrumentation, profiling,: and tracing tooling to uncover bottlenecks and guide optimizations
Architect scalable routing, batching, scheduling, memory management, and: dynamic loading mechanisms for inference workloads
Ensure reliability, reproducibility, and fault tolerance in the inference: pipelines, including A/B launches, rollback, and model versioning
What they're looking for
- Collaborate cross-functionally on Integrating with federated, distributed: inference infrastructure – orchestrate across nodes, balance load, handle communication overhead
- Drive cross-team collaboration: with platform engineers, cloud infrastructure, and security/compliance teams
- Represent the team externally through benchmarks, whitepapers, and open-source contributions
- BS/MS/PhD in Computer Science, or a related field