Staff Software Engineer - GenAI inferenceActive$191K

The opportunity

Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles.

What you'll do

  • Own and drive the architecture, design, and implementation of the inference: engine, and collaborate on model-serving stack optimized for large-scale LLMs inference

  • Partner closely with researchers to bring new model architectures or features: (sparsity, activation compression, mixture-of-experts) into the engine

  • Lead the end-to-end optimization for latency, throughput, memory efficiency,: and hardware utilization across GPUs, and accelerators

  • Define and guide standards to build and maintain instrumentation, profiling,: and tracing tooling to uncover bottlenecks and guide optimizations

  • Architect scalable routing, batching, scheduling, memory management, and: dynamic loading mechanisms for inference workloads

  • Ensure reliability, reproducibility, and fault tolerance in the inference: pipelines, including A/B launches, rollback, and model versioning

What they're looking for

  • Collaborate cross-functionally on Integrating with federated, distributed: inference infrastructure – orchestrate across nodes, balance load, handle communication overhead
  • Drive cross-team collaboration: with platform engineers, cloud infrastructure, and security/compliance teams
  • Represent the team externally through benchmarks, whitepapers, and open-source contributions
  • BS/MS/PhD in Computer Science, or a related field