San FranciscoFull TimeEngineering
Remotely
pythonllmsmlflowlangsmithbraintrustdeepevalarize phoenixopenai evals
Job Description
📋 Description
- Define quality metrics for incident diagnostics, root-cause analysis, alerting, and operational
- Build offline and online evaluation pipelines in Python and CI/CD integration.
- Lead golden datasets and regression suites using logs, metrics, events, and more.
- Create representative scenarios across technologies and environments.
- Develop regression, edge, adversarial, and incomplete-context test cases.
- Monitor AI quality and drift; translate feedback into tests and safeguards.
🎯 Requirements
- 5+ years in software, ML, or related field.
- Strong Python and production systems experience.
- Experience with AI evaluation, testing, and quality frameworks.
- Experience with multiple LLM/agent eval frameworks (LangSmith, Arize Phoenix, MLflow, OpenAI Evals
- Ability to select/integrate eval frameworks for offline/online testing and release gating.
- Knowledge of LLMs, agents, retrieval-augmented generation, and prompt engineering.
🎁 Benefits
- Competitive compensation with base salary range.
- Comprehensive health, dental, vision coverage.
- 401K with company matching; wellness programs included.
- Unlimited vacation policy and other benefits.