Mountain ViewFull TimeEngineering
Remotely
pythonvision language modelsevaluationsfttransformer modelsrl
Job Description
📋 Description
- Role focuses on evaluating AI agent systems, end-to-end evaluation pipelines, and closed-loop
- Own the evaluation pipeline end to end: data collection, loop construction, and automated hill
- Work with frontier models and open-weight vision-language models, performing supervised fine-tuning
🎯 Requirements
- Graduate degree in CS/ML or equivalent research experience; strong research judgment.
- Deep understanding of LLMs, post-training, RL, and evaluation pipelines for AI systems.
- Hands-on experience with Python and production systems; ability to run experiments end-to-end.
- Experience with agent systems, evaluation harnesses, task suites, or LLM-as-judge systems; comfort
🎁 Benefits
- Competitive base pay with annual bonus, equity, and comprehensive benefits.
- Opportunity to impact hundreds of engineers; work directly with leadership and CEO.
- Hybrid environment with direct access to compute and production data.