AI Evaluation Engineer (QA)
AppnovationTorontoFull TimeEngineering
Remotely
pythonllmqaevaluationci/cdregression
Job Description
📋 Description
- Run evals at scale across large question sets from small to millions of automated evaluations.
- Measure factual grounding and accuracy lift (before/after) statistically.
- Build metrics showing QA quality improvements over time.
- Design load and quality tests as the corpus scales.
- Define and maintain test plans, test cases, and quality gates.
- Automate regression and evaluation suites; integrate into CI/CD.
🎯 Requirements
- Bachelor’s Degree in a technical field or equivalent experience.
- 4+ years in QA / test engineering with data/ML exposure.
- Strong Python and data-science techniques for measuring grounding and answer quality.
- Experience with LLM evaluation frameworks and statistical analysis.
- Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
- Comfort working across multiple LLM providers’ outputs.
🎁 Benefits
- Equal Opportunity Employer; diversity and inclusion valued.
- Accommodations available upon request throughout recruitment.
Back to all jobs