Job Description
📋 Description Design evaluations for research judgment, hypothesis generation and testing, and long-horizon Turn real research workflows and model failures into data and evaluation flywheels. Improve model research capabilities through agent harnesses, synthetic data, RL environments, and Build and maintain safe, reliable integrations between our models and OpenAI’s research Develop research agents, experiment-orchestration systems, and sandboxed runtimes that support real Create metrics and economic models to understand RSI’s current and future effects on research 🎯 Requirements Have research or engineering experience across LLM training, model evaluations, agent systems Are a strong generalist who can move between open-ended research and practical implementation Collaborate effectively across the full stack, including systems, data, model training Are comfortable building and maintaining the data pipelines, tooling, and infrastructure needed to Are comfortable working on problems without clear definitions or established playbooks. Think rigorously about scientific quality, research taste, safety, privacy, reliability 🎁 Benefits Please note: OpenAI’s job postings include information about equity and salary; details are