The opportunity
About the Team Full Stack engineers within the Fleet Scheduling team are dedicated to building intuitive and scalable interfaces that empower researchers to efficiently manage AI workloads across some of the largest supercomputers in the world. Our focus is on developing…
What you'll do
Design and develop full-stack web applications to track, monitor, and manage: large-scale AI workloads in real time.
Collaborate with researchers and infrastructure teams to translate complex: operational needs into intuitive UIs and scalable backends.
Build data visualization tools (e.g., Gantt charts, dashboards) to provide: insights into job scheduling and resource allocation.
Optimize backend services to handle massive data throughput while ensuring: low-latency performance and high availability.
Implement frontend components that provide seamless interactions with: scheduling, storage, and compute systems.
Ensure system security, reliability, and scalability across globally: distributed supercomputing infrastructure.
What they're looking for
- Significant experience in full-stack development, with expertise in modern: frontend frameworks (React, Vue, or Angular) and backend technologies (Python, Go, or Node.js).
- Experienced in building scalable, high-performance web applications for complex distributed systems.
- Strong understanding of RESTful and GraphQL APIs, distributed databases, and: cloud infrastructure (especially Azure).
- Execution-focused with a keen eye for usability, performance, and scalability in enterprise-scale systems.