San FranciscoFull TimeEngineering
Remotely
platform engineeringllmgradingrl environmentsevaluation suitesharbor environments
Job Description
📋 Description
- Define golden sets: decompose real tasks and encode the expert quality bar.
- Build verifiers over agent trajectories and outputs, calibrated and hard to game.
- Build the eval platform that runs offline environments, task suites, and grading at scale.
- Run loss analysis over production trajectories and turn failure modes into regression tests.
- Run the optimization loop across models, prompts, skills, and harnesses.
- Own the rollout gates that decide whether an agent change ships.
🎯 Requirements
- Professional, academic, or research experience in agent engineering and evaluation, including
- Experience building evaluation suites for LLM or agent systems and familiarity with benchmarks
- Judgment about task and rubric design to measure quality improvements.
- Strong software engineering fundamentals; ability to work independently on ambiguous problems.
- Bonus: experience with Harbor environments and RL environments.
🎁 Benefits
- Up to $15k relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Generous equity grant vested over 4 years
- Free Equinox membership
- $200 monthly laundry reimbursement
Back to all jobs