The opportunity
Security is at the foundation of OpenAI’s mission to ensure that artificial general intelligence benefits all of humanity.
What you'll do
Build and evolve Detection & Response capabilities across OpenAI’s: infrastructure, products, and research environments, with an emphasis on high-signal detection and reliable operational response.
Engineer detection pipelines and tooling: develop rule lifecycle management, measurement/quality loops (coverage, precision, latency), tuning processes, and safe rollout patterns.
Automate response and investigations by building workflows that reduce toil: (triage, enrichment, containment, evidence capture) and improve time-to-understand/time-to-contain.
Partner with other Security teams and system/infrastructure owners across the: company to ensure new systems ship with the right telemetry, threat models, and response playbooks from day one.
Define D&R requirements and drive visibility across endpoints, identity,: SaaS, cloud, Kubernetes: identify telemetry/control gaps, prioritize them, and advocate for fixes with partner teams (and implement directly when it’s the fastest/most effective path).
Evaluate and respond to emergent security concerns in a frontier AI lab: environment, such as detection and response strategies for agents operating across infrastructure at scale.
What they're looking for
- Have hands-on threat detection and/or incident response experience, including: building detections, running investigations, and improving operational playbooks.
- Understand modern adversary tradecraft (TTPs) and can translate it into: practical detection strategies and response actions.
- Bring a threat modeling mindset. You can evaluate new infrastructure or: features, identify D&R implications (what could go wrong, what we’d need to see, how we’d respond), and turn that into concrete requirements for teams shipping the system.
- Have experience working in Kubernetes/containerized environments, including: building detections from cluster telemetry and understanding common failure and attack modes (workloads, nodes, control plane, networking).