The opportunity
Scale is building reliable AI systems for the world’s most important decisions. Our products help leading enterprises, governments, and frontier AI labs build, deploy, and evaluate AI systems at scale.
What you'll do
Lead the architecture, design, implementation, and operation of Scale’s core orchestration platform.
Build durable workflow infrastructure using technologies such as Temporal,: Cadence, Kubernetes, and cloud-native systems.
Define platform primitives and APIs for scheduling, retries, state: management, task execution, observability, and workflow lifecycle management.
Partner with product, infrastructure, data, and application teams to: understand workflow needs and turn them into reusable platform capabilities.
Improve the reliability, scalability, security, and developer experience of: services that run critical company workflows.
Establish technical standards and best practices for distributed workflow: development, deployment, testing, and incident response.
What they're looking for
- Drive cross-functional technical decisions and communicate platform direction: clearly to engineers and stakeholders.
- + years of full-time software engineering experience, with a focus on: backend, infrastructure, and distributed systems.
- Experience building and operating production systems with strong requirements: for reliability, availability, and scale.
- Deep familiarity with workflow orchestration platforms such as Temporal,: Cadence, AWS Step Functions, Kubernetes, or similar systems.