Senior Machine Learning Engineer, Agent Eval Platform
ServiceNowSanta ClaraFull TimeEngineering
Remotely
pythonllmevaluationrlhfcalibrationreward model
Job Description
📋 Description
- Build the judgement layer of our agent evaluation platform: rubrics, judges, and calibration
- Develop deterministic validators for checkable world state and LLM judges for fuzzy parts.
- Implement scoring that reports confidence and handles uncertain judgments with human routing.
- Establish calibration loops with human-labeled trajectories and annotation teams.
- Fine-tune small judges and guard against blind spots in simulations vs. production.
- Design a scalable substrate for future RL efforts with versioned scenarios and reward signals.
🎯 Requirements
- 5+ years in applied ML, data science, or ML-adjacent engineering with shipped impact.
- Experience turning subjective human judgement into a measurable signal.
- Strong Python and ability to ship production-grade code.
- Ability to communicate complex problems clearly and influence engineers.
- High ownership, startup-paced shipping mindset, and comfort with ambiguity.
- Experience in at least 3 of the following: LLM evaluation design, human annotation programs, online
🎁 Benefits
- Work with an AI-driven agent platform within a fast-growing environment.
- Impactful role at the intersection of ML research and production systems.
- Collaborative team with opportunities to influence product direction.
Back to all jobs