The opportunity
OpenAI’s People team hires, engages, and retains world-class talent to safely build and deploy AGI that benefits all of humanity. The People Analytics team helps leaders make rigorous, evidence-based talent decisions and ensures that the systems supporting those decisions are valid, reliable, fair, and accountable.
What you'll do
Define and lead fairness and bias-testing strategies for AI-assisted People: processes, models, agents, and decision-support systems from development through deployment and ongoing monitoring.
Design rigorous algorithmic audits and validation studies, including: adverse-impact analysis, subgroup and intersectional evaluation, error-rate analysis, calibration, measurement invariance, reliability, criterion-related validity, and sensitivity testing.
Identify the appropriate fairness criteria for each use case, evaluate: tradeoffs among competing definitions of fairness, and clearly document the assumptions, limitations, and residual risks of each approach.
Evaluate end-to-end human-AI decision systems, including model outputs, user: behavior, human overrides, escalation pathways, and whether AI assistance changes the quality, consistency, or equity of decisions.
Develop evaluation approaches for generative and agentic AI, including: test-set design, counterfactual testing, behavioral evaluation, human-rating studies, robustness testing, and analysis of disparate performance across populations and contexts.
Investigate the sources of observed disparities, including data: representation, label and measurement bias, proxy variables, model design, decision thresholds, workflow design, and differential adoption or usage.
What they're looking for
- Experience conducting fairness assessments, algorithmic audits, model-risk: reviews, adverse-impact analyses, or validation studies in employment or another high-impact domain.
- Familiarity with fairness and model-evaluation tools such as Fairlearn, AI: Fairness 360, responsible-AI evaluation frameworks, explainability methods, or comparable internal tooling.
- Experience evaluating large language models, generative AI systems, safety: classifiers, or agentic workflows, including behavioral testing and human evaluation.
- Experience with employment selection, talent assessment, psychometrics,: organizational research, or the validation of hiring, performance, promotion, or workforce decisions.
- Familiarity with responsible-AI frameworks and emerging requirements related: to automated employment decision systems, algorithmic auditing, data privacy, and AI governance.
- Experience creating model cards, dataset documentation, fairness scorecards,: audit reports, monitoring plans, or other review artifacts for high-impact systems.
- Advanced degree in Quantitative Psychology, Computer Science, Statistics,: Economics, Data Science, Behavioral Science, or a related quantitative field; PhD preferred but not required.