Sr. Evaluation Engineer
LogicMonitorJob Description
📋 Description Define quality metrics for incident diagnostics, root-cause analysis, and alert correlation. Build offline and online evaluation pipelines in Python; CI/CD integration. Lead golden datasets and regression suites across metrics, logs, and incidents. Create customer-specific scenarios across technologies and failure modes. Use human-authored and AI-assisted methods to generate regression and edge test cases. Monitor AI quality and drift; establish evaluation-driven development; mentor engineers. 🎯 Requirements 5+ years in software engineering, ML, AI, or related field. Strong Python engineering and production systems experience. Hands-on AI evaluation, experimentation, testing, and quality frameworks. Experience with LangSmith, Arize Phoenix, MLflow, or similar evaluation tools. Ability to select/integrate evaluation frameworks for offline testing, online monitoring, regression analysis, and release gating. Strong understanding of LLMs, agents, retrieval-augmented generation, prompt engineering, tool calling. Experience evaluating non-deterministic, multi-step, or multi-agent AI systems. Experience with regression testing, CI/CD, production monitoring, drift detection, and failure analysis. 🎁 Benefits Great Place To Work certification. BuiltIn's Best Places to Work for the seventh year. Equal opportunity employer with an inclusive culture.