The opportunity
The Technical Support team is responsible for ensuring that developers and enterprises can reliably build mission critical solutions using OpenAI models. We provide technical guidance, resolve complex issues and support customers in maximizing value and adoption from deploying our highly-capable models.
What you'll do
Be among the foremost technical and troubleshooting experts for our API: platform at OpenAI. You are the last line of defense before the core Engineering team.
Proactively identify and implement opportunities to scale support operations: by leveraging automation and advancements in AI technologies. Contribute to shaping the future of technical support in an AI-driven era.
Configure and use advanced monitoring and alerting workflows to proactively: detect customer impacting issues in real time.
In partnership with engineering, contribute to reliability reviews and: preparedness for new features, launches, or strategic customer requirement updates. Ensure that operational readiness (monitoring, alerting, and fallback plans) is in place for any such changes.
Design and refine incident response processes and documentation across: strategic customers, engineering and support teams.
Analyze operational metrics and incident RCAs to identify areas for: improvement. Proactively recommend and implement enhancements to monitoring dashboards, alert configurations, and support workflows.
What they're looking for
- Provide support coverage during holidays and weekends based on business needs.
- Have a Bachelor’s degree in Computer Science or a related field. A strong: software engineering foundation is important for this role’s success.
- Have 8+ years of experience in technical operations roles such as SRE/NOC,: designing monitoring systems and resolving production issues in fast-paced and mission-critical environments. A strong track record of troubleshooting complex technical problems at the systems level.
- Have deep familiarity with modern monitoring, alerting, and observability: practices. Hands‑on experience setting up or managing metrics, logging, and tracing for distributed systems (e.g., understanding of SLIs/SLOs, alert tuning, dashboard creation).