Job Description
📋 Description Post-Training team adapts foundation models to real-world performance and alignment. Develop/evaluate supervised fine-tuning, preference optimization (DPO/RLHF/RLAIF), continual Bridge raw model capability with trustworthy, contextually aligned system behavior. Align large models with human and system-level objectives across industries. Explore trade-offs between generalization, data efficiency, robustness, capability and Leverage AI-native tools to accelerate research workflows. 🎯 Requirements Deep understanding of post-training techniques: supervised fine-tuning, RLHF/DPO, LoRA/PEFT Experience adapting frontier models to specialized domains via data curation or reward modeling. Strong ability to build prototypes and run experiments to prove effectiveness to enterprise Experience with compound AI systems, agentic collaboration, ensembling, ReAct, graph-of-thoughts Proven track record of research results (publications, notable work). Daily use of AI tools (ChatGPT, Cursor, Perplexity) to accelerate workflows. 🎁 Benefits Salary: $150K – $250K base, plus meaningful equity and comprehensive benefits. 100% coverage of medical, dental, and vision for dependents. Flexible time off and retirement benefits, including 401(k) and HSA/FSA options. Wellness programs and in-office meals; access to state-of-the-art AI models/tools. Opportunity to own high-impact projects across top enterprises. Hybrid work model with in-office days; supportive, mission-driven culture.