Senior HPC Systems Validation EngineerActive$255K–$340K

The opportunity

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.

What you'll do

  • Own system integration validation for new HPC AI/ML, general purpose compute,: storage, and network hardware platforms throughout hardware NPI and deployment readiness.

  • Development and execution of functional, stress, reliability, error-injection: and benchmark tests on new hardware systems. Enable automated tools for doing scale testing.

  • Collaborate with HPC deployment and operation teams, fleet engineering teams,: infrastructure engineering and PMO teams to facilitate tooling automation and L10/L11/L12 level benchmarking throughout hardware NPI and at new hardware deployment.

  • Manage the firmware and software compatibility, track and recommend firmware: and software releases during NPI, support HPC operation team on firmware and software maintenance after production release.

  • Support RMA team on critical hardware issues’ triage and Root-Cause-Analysis: after deployment; own fleet hardware issues’ failure correlation, trend analysis, close the loop with Platform Hardware Engineer and Quality team to improve product reliability.

  • Own hands-on lab setup and test infrastructure for repeatable hardware: evaluation, partnering with Platform Hardware Engineer, Data Center Engineering, and Cluster Network Design to ensure new platforms can be brought up, instrumented, stressed, benchmarked, and debugged consistently.

What they're looking for

  • Manage hardware vendors regarding hardware validation and manufacturing: testing; support new vendor evaluation and Quarterly Business Review on hardware vendors.
  • Drive alternate component qualification from technical evaluation through PLM: enablement, develop and execute validation plans for key commodities such as memory, SSDs, NICs, PSUs, PDUs, fans, cables, optics, and transceivers.
  • years of experience in one or many of the following areas: hardware integration validation, fleet hardware reliability engineering, component qualification, performance validation for HPC, data center, or cloud infrastructure products.
  • Possess deep knowledge in system integration testing and performance: benchmarking at one or many of the L10, L11 and L12 levels.