San FranciscoFull TimeEngineering
Remotely
llmevaluationrlhfsftmultimodalreward modelingbenchmark
Job Description
📋 Description
- Analyze model behavior to identify and diagnose failure modes in frontier LLMs and Agents.
- Design and build benchmarks and evaluation methods for text and multimodal data.
- Apply post-training expertise (SFT, RLHF, reward modeling) to address observed failures.
- Publish research findings in top AI conferences and journals.
🎯 Requirements
- Ph.D. or Master's in CS, ML, AI, or related field.
- Deep understanding of DL, RL, and large-scale model fine-tuning.
- Experience with post-training techniques (RLHF, preference modeling, or instruction tuning) and LLM
- Strong written and verbal communication; published ML research.
- Customer-facing experience.
🎁 Benefits
- Compensation includes base salary, equity, and benefits.
- Health, dental, vision coverage; retirement benefits; learning stipend; generous PTO.
- Possible commuter stipend and other benefits as applicable.
Back to all jobs