Member of Technical Staff (AI Inference Engineer)Active

LondonTechnology

The opportunity

We are looking for an AI Inference Engineer to join our growing team. We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets.

What they're looking for

  • You understand modern LLM architectures and are able to bring them up reliably in a production environment.
  • You've built and operated production distributed systems under real load - ideally performance-critical ones.
  • Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels.
  • You own problems end-to-end. You can read a research paper on Monday, write a: kernel on Wednesday, and debug a production incident on Friday.
Member of Technical Staff (AI Inference Engineer) at Perplexity | Role Match