Job Description
📋 Description Research multimodal safety for text, vision, and audio; connect perception to safe model behavior. Build training/evaluation methods for vision-language models; post-training, safety evals, and Collaborate with internal teams to translate research into safer multimodal experiences. Role based in San Francisco with a hybrid work model (3 days in office/week) and relocation 🎯 Requirements Strong track record in multimodal models; depth in vision-language, video understanding, image Understanding of end-to-end multimodal systems: encoders, fusion, cross-modal reasoning, scaling Experience with post-training improvements (SFT, RL) and robust evaluation; data curation and Ability to form hypotheses, run decisive experiments, diagnose failures, and translate findings Excellent safety judgment; strong research and engineering judgment for open-ended safety problems. 🎁 Benefits Relocation assistance OpenAI-specific benefits and equal opportunity employer commitment Competitive compensation with equity options