Member of Technical Staff, Performance OptimizationActive$175K–$220K

The opportunity

Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed,…

What you'll do

  • Optimize system and GPU performance for high-throughput AI workloads across training and inference

  • Analyze and improve latency, throughput, memory usage, and compute efficiency

  • Profile system performance to detect and resolve GPU: and kernel-level bottlenecks

  • Implement low-level optimizations using CUDA, Triton, and other performance tooling

  • Drive improvements in execution speed and resource utilization for: large-scale model workloads (LLMs, VLMs, and video models)

  • Collaborate with ML researchers to co-design and tune model architectures for hardware efficiency

What they're looking for

  • Improve support for mixed precision, quantization, and model graph optimization
  • Build and maintain performance benchmarking and monitoring infrastructure
  • Scale inference and training systems across multi-GPU, multi-node environments
  • Evaluate and integrate optimizations for emerging hardware accelerators and specialized runtimes
Member of Technical Staff, Performance Optimization at Fireworks AI | Role Match