Job Description
📋 Description Own the scientific validity of frontier preparedness evaluations Design new evals grounded in real threat models (CBRN, cyber, frontier risks) Define datasets, graders, rubrics, and threshold guidance Produce auditable artifacts (evaluation cards, capability reports, system-card inputs) Build scalable systems to support evaluations Collaborate with leadership to trust high-stakes launches 🎯 Requirements Experience in ML research engineering or ML observability Ability to build evaluations for frontier AI models Strong data handling: datasets, graders, rubrics, and thresholding Red-teaming mindset and capability with risk assessment Excellent cross-functional communication 🎁 Benefits Equity options and market-competitive salary OpenAI benefits and growth opportunities