AI QA & Evaluation Engineer
ElasticRemotely
pythontypescriptgithubterraformaivertex ailangsmithazure openai
Job Description
📋 Description
- Be a primary contributor to our AI strategy, helping validate and test AI infrastructure, custom
- Test strategy and execution for AI/ML systems, including accuracy, bias, robustness, and regression
- Design and implement self-contained evaluation tasks, prompts, and grading rubrics for GenAI
- Automate validation suites for agentic systems and ML model CI/CD pipelines.
- Observe AI agent behaviors and document performance, focusing on reliability and data grounding.
- Collaborate across IT, Engineering, Data, and Operations to build robust testing frameworks.
🎯 Requirements
- Proficiency in Python, TypeScript, or other programming languages used in AI and test automation.
- Experience with rubric-based evaluation, scoring frameworks, or structured grading methods.
- Experience with LLM evaluation frameworks and benchmarking tools (e.g., LangSmith, Confident AI).
- Direct experience with GenAI stacks (Retrieval Augmented Generation) and related tooling (Azure
- Strong written communication to document observations and provide actionable feedback.
- Knowledge of DevOps/CI/CD tools (GitHub, Terraform) and data governance considerations.
🎁 Benefits
- Experience with cloud platforms (Azure, GCP, AWS).
- Competitive compensation and inclusive, distributed work culture.
- Flexible locations and schedules for many roles, with generous vacation time.
- Support for professional development and ongoing learning in AI governance and evaluation.
Back to all jobs