AI Evaluation Engineer (QA)
AppnovationRemotely
pythondata analysistest automationqallm evaluationregression testingci/cd
Job Description
📋 Description
- QA / AI Evaluation Engineer role in a forward-leaning team.
- Run evaluations at scale from small human-UAT batches to millions of automated evals.
- Measure factual grounding and accuracy lift; build metrics framework.
- Design load and quality tests as corpus scales; automate regression and evaluation suites
- Report quality metrics to technical and non-technical stakeholders; collaborate with engineering to
🎯 Requirements
- Bachelor’s Degree in a technical field or equivalent experience.
- 4+ years in QA / test engineering with exposure to data/ML systems.
- Strong Python and data-science techniques for measuring factual grounding and answer quality.
- Experience with LLM evaluation frameworks and statistical analysis.
- Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
- Comfort working across multiple LLM providers’ outputs; test automation frameworks and scripting.
🎁 Benefits
Back to all jobs