Member of Technical Staff, Evals & Post-Training ProductActive$200K–$350K
The opportunity
Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed,…
What you'll do
Scale Eval Infrastructure: Take ownership of our internal eval setup and evolve it for the future. You will design systems to eliminate single-host coordination bottlenecks, resolve log-syncing latency, and build seamless programmatic access.
Benchmark Obsession & Reproduction: Act as a "metrics obsessive." Track state-of-the-art (SOTA) benchmarks, read the latest research papers, dig deep into data discrepancies, and insist on rigorously reproducing published results internally.
Pioneer New Benchmarks: Design and build entirely new benchmarks to measure model performance on complex, emerging, or domain-specific use cases.
Own Fine-Tuning Product Experiences: Build and improve user-facing workflows for post-training, including fine-tuning experiences across SFT, RFT, and related model-improvement capabilities.
Work Closely With Users: Partner with customers and internal stakeholders to understand evaluation and fine-tuning needs, triage issues, and convert bespoke workflows into productized, reusable solutions.
Experience: 1–7 years of software engineering or data science experience (We: are hiring at multiple levels for this role).
What they're looking for
- Experience: 1–7 years of software engineering or data science experience (We: are hiring at multiple levels for this role).
- Strong System Design Skills: You know how to architect scalable, programmatic systems and transition legacy setups into robust infrastructure.
- Sandbox Infrastructure: Hands-on experience building or working with sandbox environments and sandbox infrastructure for secure code execution and testing.
- Analytical & Data Science Mindset: You possess a deep understanding of LLM evaluations, how to design them, and how to use the results to guide model improvement. You are meticulous about data and metrics.
- Understanding of the GenAI Lifecycle: You understand the end-to-end workflow—from prompting a base model to curating a dataset, fine-tuning, and productionizing agents.
- Experience: 3+ years of software engineering or applied data science experience.
- Frameworks & Orchestration: Experience working with the Harbor framework or similar container registry and orchestration tools.
- Public Writing & Analysis: A strong interest in discovering where different models excel and fall short, with a desire to write up and publish these insights publicly (e.g., technical blogs, whitepapers).