The opportunity
Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads.
What you'll do
Bring current inference techniques into production and refine them.
Design and optimize serving architectures, including prefill and decode: disaggregation, request routing, and related approaches.
Work down into the serving stack, from frameworks like vLLM and SGLang to the: CUDA kernels underneath, profiling and running in-depth analysis to find and fix performance problems.
Adapt and scale optimization methods across many kinds of ML models, with an: emphasis on large language models.
Profile and tune deployments against clear targets for latency, throughput,: and cost, and keep them dependable under real traffic.
Tailor deployments to each customer's models and constraints, partnering with: their engineering teams to move a workload from an early proof of concept through to a live, well-monitored production service.
What they're looking for
- Build and support the software and product features around the inference: stack in a production setting, using one or more general-purpose languages, with Python preferred given how central it is to ML work.
- Experiment quickly: take fuzzy goals, shape them into clear specs and focused proofs of concept, run fast experiments to find what works, and ship well-tested results without delay.
- Own delivery end to end, from the first experiment through to the: optimization running in production, keeping the underlying performance goals, clear specs, and follow-through front of mind, and drafting features and product requirement documents together with other engineering and product teams.
- Work through ambiguity and make sound calls on tradeoffs and tooling,: steering away from complexity that is not needed.